Disease prediction method, device, equipment and storage medium

By generating the target disease characteristic spectrum, the problem of low accuracy in the existing disease prediction methods is solved, and efficient and accurate disease prediction is achieved.

CN115116618BActive Publication Date: 2025-08-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210614057.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-08-12
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing disease prediction methods are greatly affected by subjective factors and have low accuracy, so it is impossible to effectively determine the disease patients suffer from.

Method used

By determining the initial disease characteristic spectrum of the disease information set, fusing each initial disease characteristic spectrum to generate the target disease characteristic spectrum, and predicting the patient's target disease based on the target disease characteristic spectrum.

Benefits of technology

It improves the accuracy and efficiency of disease prediction, is highly applicable, and can accurately determine the target disease that the patient suffers from.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116618B_ABST
    Figure CN115116618B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a disease prediction method, apparatus, device, and storage medium applicable to fields such as cloud technology, artificial intelligence, and blockchain. The method comprises: determining at least one disease information set, each disease information set including at least one disease suffered by a patient and the symptoms present; for each disease information set, determining an initial disease characteristic spectrum corresponding to the disease information set; determining a target disease characteristic spectrum based on each initial disease characteristic spectrum; determining at least one target symptom of the patient to be predicted, and determining the target disease suffered by the patient to be predicted based on each target symptom and the target disease characteristic spectrum. The embodiments of the present application can accurately and efficiently determine the target disease suffered by the patient to be predicted, and have high applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a disease prediction method, apparatus, device, and storage medium. Background Art

[0002] Disease refers to the abnormal life activity process caused by the disorder of the body's self-stabilizing regulation under certain causes. Different diseases can manifest different symptoms. Disease prediction refers to inferring the disease a patient suffers from based on the symptoms he or she suffers from.

[0003] In the existing technology, researchers usually design a scoring table based on clinical experience, assign a quantitative score to each symptom feature, and further calculate the total score or normalized value of each symptom to predict the disease suffered by the patient. For example, Figure 1 A matching score table is used to predict venous thromboembolism. When the total score of a patient's symptoms is higher than 6 points, the patient is considered to have possible venous thromboembolism.

[0004] The prediction results of existing technologies are greatly affected by subjective factors. Often, due to the lack of rigorous data verification and fitting, the final prediction results are significantly different from the actual results and the accuracy is low. Summary of the Invention

[0005] The embodiments of the present application provide a disease prediction method, apparatus, device, and storage medium, which can improve the accuracy and efficiency of disease prediction and have high applicability.

[0006] In one aspect, an embodiment of the present application provides a disease prediction method, the method comprising:

[0007] Determine at least one disease information set, each of which includes at least one disease suffered by a patient and the symptoms that occur;

[0008] For each of the above disease information sets, determining an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set;

[0009] Determine a target disease characteristic spectrum based on each of the initial disease characteristic spectra, wherein the target disease characteristic spectrum represents a target probability of each preset symptom occurring in each preset disease, each of the target probabilities being determined based on at least one first probability of the preset disease corresponding to the target probability occurring in each of the initial disease characteristic spectra, each of the preset diseases being one of the diseases corresponding to each of the disease information sets, and each of the preset symptoms being one of the symptoms corresponding to each of the disease information sets;

[0010] Determine at least one target symptom of the patient to be predicted, and based on the target symptoms and the target disease characteristic spectrum, determine the target disease suffered by the patient to be predicted.

[0011] On the other hand, an embodiment of the present application provides a disease prediction device, which includes:

[0012] An information determination module, configured to determine at least one disease information set, each of which includes at least one disease suffered by a patient and the symptoms that occur;

[0013] a probability determination module, configured to determine, for each of the disease information sets, an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set;

[0014] a probability fusion module for determining a target disease characteristic spectrum based on each of the initial disease characteristic spectra, wherein the target disease characteristic spectrum represents a target probability of each preset symptom occurring in each preset disease, each target probability being determined based on at least one first probability of the preset disease corresponding to the target probability occurring in each of the initial disease characteristic spectra, each of the preset diseases being one of the diseases corresponding to each of the disease information sets, and each of the preset symptoms being one of the symptoms corresponding to each of the disease information sets;

[0015] The disease prediction module is used to determine at least one target symptom of the patient to be predicted, and based on the above target symptoms and the above target disease characteristic spectrum, determine the target disease suffered by the above patient to be predicted.

[0016] On the other hand, an embodiment of the present application provides an electronic device, including a processor and a memory, wherein the processor and the memory are connected to each other;

[0017] The memory is used to store computer programs;

[0018] The above-mentioned processor is configured to execute the disease prediction method provided in the embodiment of the present application when calling the above-mentioned computer program.

[0019] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the disease prediction method provided by the embodiment of the present application.

[0020] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program. When the above computer program is executed by a processor, it implements the disease prediction method provided by the embodiment of the present application.

[0021] In an embodiment of the present application, by determining the initial disease characteristic spectrum corresponding to each disease information set, the first probability of each symptom of each disease in each disease information set can be preliminarily determined, and then after fusing the initial characteristic spectra, the target disease characteristic spectrum that characterizes the target probability of each preset symptom of each preset disease can be accurately obtained, thereby having high accuracy and efficiency when determining the target disease of the patient to be predicted based on the target disease characteristic spectrum, and high applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A matching score table for predicting venous thromboembolism in the prior art;

[0024] Figure 2 Schematic diagram of the disease prediction method provided in the embodiment of the present application;

[0025] Figure 3 This is a schematic diagram of outpatient records provided by an embodiment of the present application;

[0026] Figure 4 is a schematic diagram of a symptom matrix provided in an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of the disease matrix provided in the embodiments of the present application;

[0028] Figure 6 is a schematic diagram of the disease characteristic spectrum provided in the examples of this application;

[0029] Figure 7 is a schematic diagram of the symptom relationship spectrum provided in the examples of the present application;

[0030] Figure 8 is a schematic structural diagram of a disease prediction device provided in an embodiment of the present application;

[0031] Figure 9 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0033] The disease prediction method provided in the embodiments of the present application can be applied to a disease control center or an infectious disease prevention system, etc., to accurately determine the disease suffered by the patient and improve the accuracy and efficiency of disease prediction.

[0034] See also Figure 2 , Figure 2 Schematic diagram of the process of disease prediction provided by the embodiment of the present application. Figure 2 As shown, the disease prediction method provided in the embodiment of the present application may include the following steps:

[0035] Step S21: Determine at least one disease information set.

[0036] In some feasible implementations, each disease information set includes at least one disease and symptoms suffered by a patient. Different disease information sets correspond to different information sources, including but not limited to published literature, disease statistics, and patient-provided medical information or outpatient records. The specific information sources can be determined based on the actual application scenario requirements and are not limited here.

[0037] The acquisition of information related to any patient's disease and symptoms, and its application to the disease prediction methods provided in the embodiments of this application, requires the patient's permission or consent. The acquisition of literature, statistical data, and medical records related to disease and symptom information also requires the permission or authorization of the information owner. The collection, use, and processing of information related to diseases and symptoms must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0038] like Figure 3 As shown, Figure 3 This is a schematic diagram of outpatient records provided by the embodiment of this application. Figure 3 The outpatient records shown indicate that the patient developed a fever with no apparent cause one week ago, reaching a maximum temperature of 37.6°C. He also experienced headache, fatigue, and body aches, but no runny nose or nasal congestion, a slight cough, no sputum, chest tightness or pain, nausea or vomiting, abdominal pain or diarrhea, frequent urination, urgency, or pain, or lower back pain. Self-medication (specific medications are unknown) failed to alleviate the symptoms, and the patient had visited a fever clinic. By performing named entity recognition, relation extraction, and text matching on this outpatient record, it was determined that the patient suffered from an acute upper respiratory tract infection, and that the symptoms included fever, fatigue, headache, and a dry cough.

[0039] Different patients in the same disease information set may suffer from the same disease or different diseases, and different patients may have different symptoms when suffering from the same disease.

[0040] Among them, the above-mentioned diseases may include common diseases such as acute upper respiratory tract infection (cold), acute bronchitis, acute gastroenteritis, etc., and may also include rare diseases such as amyotrophic lateral sclerosis and other diseases with lower incidence rates. The specific diseases can be determined based on the actual application scenario requirements and are not restricted here.

[0041] Among them, the symptoms of any patient may include the patient's physical symptoms, such as fever, headache, runny nose and cough, or may also include related symptoms discovered by imaging equipment such as ground-glass shadows in the lungs, or may also include abnormal indicators obtained through laboratory tests, such as decreased white blood cell count, abnormal platelet count, etc., or may also include the patient's characteristics, such as age, allergy history, obesity rate, etc.

[0042] The above descriptions of the symptoms are only examples and can be determined based on the actual application scenario requirements and are not limited here.

[0043] Step S22: For each disease information set, determine the initial disease characteristic spectrum corresponding to the disease information set.

[0044] In some feasible embodiments, the initial disease characteristic spectrum corresponding to each disease information set represents the probability of each symptom in the disease information set occurring in each disease in the disease information set (hereinafter referred to as the first probability for the convenience of description), that is, it represents the first probability of each symptom occurring when the patient suffers from any disease in the disease information set.

[0045] Specifically, for each disease and each symptom in each disease information set, a first number of patients suffering from the disease and a second number of patients suffering from the disease and experiencing the symptom in the disease information set may be determined.

[0046] For example, a disease information set includes the disease "cold" and the symptom "cough", and a first number of patients suffering from a cold and a second number of patients suffering from and experiencing a cough symptom in the disease information set can be determined.

[0047] Furthermore, for each disease and each symptom in each disease information set, the ratio of the second number of patients suffering from the disease and experiencing the symptom to the first number of patients suffering from the disease can be determined as the first probability of the disease experiencing the symptom.

[0048] For each disease information set, the disease matrix and symptom matrix corresponding to the disease information set can also be determined, and then the first probability of any symptom of any disease occurring can be determined through the disease matrix and symptom matrix.

[0049] For each disease information set, the symptom matrix corresponding to the disease information set can be determined based on the symptoms of each patient in the disease information set. v *N f dimensional matrix, X ih Indicates the value of the hth feature of the i-th patient, which is used to indicate whether the i-th patient has the h-th symptom. v is the total number of patients in the disease information set, N f is the number of symptom types in the disease information set.

[0050] like Figure 4 As shown, Figure 4 This is a schematic diagram of the symptom matrix provided by the embodiment of the present application. After converting the patient's symptoms into 0 / 1 discrete feature vectors, when the vector value of an element in the symptom matrix is 1, it indicates that the corresponding patient has the corresponding symptom, and when the vector value is 0, it indicates that the corresponding patient does not have the corresponding symptom. For example, based on Figure 4 From the vector values in the first column and first row of the symptom matrix shown, it can be seen that patient 1 developed coughing symptoms when he was ill.

[0051] For each disease information set, the disease matrix corresponding to the disease information set can be determined based on the diseases suffered by each patient in the disease information set. v *N d dimensional matrix, Y ik The value of the kth disease of the i-th patient is used to indicate whether the i-th patient has the k-th disease. v is the total number of patients in the disease information set, N d is the number of disease types in the disease information set.

[0052] Figure 5 As shown, Figure 5 This is a schematic diagram of the disease matrix provided by the embodiment of the present application. After converting the manifestation of each disease of the patient into a 0 / 1 discrete feature vector, when the vector value of an element in the disease matrix is 1, it means that the corresponding patient has the corresponding disease, and when the vector value is 0, it means that the corresponding patient does not have the corresponding disease. For example, based on Figure 5 From the vector value in the third column and second row of the symptom matrix shown, we can see that patient 2 suffers from diabetes.

[0053] For the disease information set r, the first probability of disease k having symptom h corresponding to the disease information set is This can be determined by:

[0054]

[0055] Among them, Y ik Indicates whether the i-th patient in the disease matrix has disease k, X ih Indicates whether the i-th patient in the symptom matrix has symptom h, N v is the total number of patients in the disease information set.

[0056] When the vector value of an element in the symptom matrix is 1, it indicates that the corresponding patient has the corresponding symptom, and when the vector value is 0, it indicates that the corresponding patient does not have the corresponding symptom; and when the vector value of an element in the disease matrix is 1, it indicates that the corresponding patient has the corresponding disease, and when the vector value is 0, it indicates that the corresponding patient does not have the corresponding disease. represents the first number of patients with disease k, represents the second number of patients with disease k and symptom h.

[0057] Furthermore, for each disease information set, after determining the first probability of each symptom in the disease information set occurring for each disease in the disease information set, an initial disease characteristic spectrum corresponding to the disease information can be constructed based on each first probability.

[0058] like Figure 6 As shown, Figure 6 Schematic diagram of the disease characteristic spectrum provided in the embodiment of this application. Figure 6 As shown, the disease characteristic spectrum can be regarded as a matrix, and the vector value of an element in the matrix represents the probability of the corresponding disease appearing. Figure 6 From the vector value in the second row of the third column in the symptom matrix shown, it can be seen that the probability of acute bronchitis having a symptom of expectoration is 0.8, that is, the probability of a patient with acute bronchitis having a symptom of expectoration is 0.8.

[0059] Step S23: Determine the target disease characteristic spectrum based on each initial disease characteristic spectrum.

[0060] In some feasible implementations, since the diseases suffered by patients and the symptoms that occur in different disease information sets are somewhat random, the probability of each item in the corresponding initial disease characteristic spectrum may differ from the actual probability. Therefore, in order to improve the first probability of any symptom of any disease to be more accurate, the initial disease characteristic spectra corresponding to each disease information set can be fused to obtain the target disease characteristic spectrum.

[0061] Among them, the target disease characteristic spectrum represents the target probability of each preset symptom occurring in each preset disease, each preset disease is one of all diseases corresponding to each disease information set, and each preset symptom is one of all symptoms corresponding to each disease information set.

[0062] For example, the preset diseases may be diseases that are commonly present in each of the initial disease profiles, and the preset symptoms may be symptoms that are commonly present in each of the initial disease profiles. Alternatively, the preset diseases may be diseases that appear with a frequency greater than a certain threshold in each of the initial disease profiles, and the preset symptoms may be all symptoms corresponding to each of the preset diseases in each of the initial disease profiles. Alternatively, the preset diseases may be all diseases corresponding to each of the initial disease profiles, and the preset symptoms may be all symptoms corresponding to each of the initial disease profiles.

[0063] It should be noted that the above-mentioned methods for determining preset diseases and preset symptoms are only examples and can be determined based on the actual application scenario requirements and are not limited here.

[0064] Among them, for each target probability in the target disease characteristic spectrum, the target probability is determined based on the respective first probabilities of the preset disease corresponding to the target probability in all initial disease characteristic spectra exhibiting the preset symptoms corresponding to the target probability.

[0065] Specifically, for each preset disease and each preset symptom, each first probability of the preset disease appearing with the preset symptom can be determined from all initial disease characteristic spectra, and then the average probability of each determined first probability can be determined as the target probability of the preset disease appearing with the preset symptom. Alternatively, a weight can be determined for each initial disease characteristic spectrum, and the weight of each initial disease characteristic spectrum represents the credibility or data validity of the initial disease characteristic spectrum, and can also represent the validity of the corresponding disease information set. Then, each first probability of the preset disease appearing with the preset symptom can be multiplied by the corresponding weight, and the multiplication results can be further added together to obtain the target probability of the preset disease appearing with the preset symptom.

[0066] Optionally, for each preset disease and each preset symptom, a disease characteristic spectrum with a first probability of including the preset disease in each initial disease characteristic spectrum (hereinafter referred to as the first disease characteristic spectrum for the convenience of description) can be determined, that is, the first disease characteristic spectrum including relevant information of the preset disease is determined from each initial disease characteristic spectrum.

[0067] Furthermore, the third number of each first disease characteristic spectrum and the first probability of the preset disease appearing with the preset symptom in each first disease characteristic spectrum are determined, and the fusion probability of the preset disease appearing with the preset symptom is determined.

[0068] Specifically, the probability sum of the first probabilities of the preset disease appearing with the preset symptom in each first disease characteristic spectrum can be used, and then the ratio of the probability sum to the third number of the first disease characteristic spectrum can be determined as the fusion probability of the preset disease appearing with the preset symptom. Specifically, the determination can be based on the following method:

[0069]

[0070] in Indicates the fusion probability of the preset disease k and the preset symptom j. 1(s,k) is an indicative function. A value of 1 indicates that the disease signature spectrum s includes relevant information of disease k, that is, it is used to represent the first disease signature spectrum that includes the first probability of disease k. A value of 0 indicates that the disease signature spectrum s is not the first disease signature spectrum. It represents the first probability of disease k having symptom j in disease feature spectrum s.

[0071] Among them, N s represents the number of disease characteristic spectra, A third quantity can be used to represent the first disease profile.

[0072] Alternatively, the weight of each first disease characteristic spectrum can be determined, and then the first probability of the preset symptom of the preset disease in each first disease characteristic spectrum can be multiplied by the weight of the corresponding first disease characteristic spectrum, and the multiplication results corresponding to each first disease characteristic spectrum can be added together to obtain an addition result. The ratio of the addition result to the third number of the first disease characteristic spectrum is determined as the fusion probability of the preset symptom of the preset disease. Specifically, it can be determined based on the following method:

[0073]

[0074] Among them, w s Represents the weight of the disease feature spectrum s.

[0075] Furthermore, based on the above method, the fusion probability of each preset symptom occurring in each preset disease can be determined, and then the characteristic spectrum of the target disease can be determined based on each fusion probability.

[0076] In some feasible implementations, when determining the target disease characteristic spectrum based on the fusion probability of each preset symptom occurring in each preset disease, for each preset disease and each preset symptom, the fusion probability of the preset disease occurring in the preset symptom can be determined as the target probability of the preset disease occurring in the preset symptom, and then the target disease characteristic spectrum is determined based on the target probability of each preset symptom occurring in each preset disease, which can be specifically Figure 6 The form shown will not be described in detail here.

[0077] In some feasible implementations, for each preset symptom, one preset symptom may be a sub-symptom of another preset symptom, for example, the preset symptom "sputum cough" is a sub-symptom of "cough." Furthermore, for each preset disease and each preset symptom, the probability of the preset symptom occurring in the preset disease must be greater than the probability of the sub-symptom of the preset symptom occurring in the preset disease.

[0078] Based on this, to further improve the accuracy of the target disease signature spectrum, we can first determine the symptom relationship spectrum corresponding to each preset symptom. This symptom relationship spectrum represents the subordinate relationship between the preset symptoms. Among them, a node in the symptom relationship spectrum represents a preset symptom, and the preset symptoms represented by the child nodes of a node in the symptom relationship spectrum are sub-symptoms of the preset symptom represented by the node.

[0079] Among them, the symptom relationship spectrum corresponding to each preset symptom can be represented by a directed acyclic graph.

[0080] See also Figure 7 , Figure 7 It is a schematic diagram of the symptom relationship spectrum provided by the embodiment of the present application. Assuming that there are preset symptoms of "cough", "sputum", "dry cough", "hemoptysis" and "blood in sputum", it can be determined that "sputum", "dry cough" and "hemoptysis" are sub-symptoms of "cough", and "blood in sputum" is a sub-symptom of "sputum" and "hemoptysis". Based on the subordinate relationship of the above preset symptoms, it can be obtained Figure 7 Symptom relationship spectrum shown.

[0081] Among them, "cough" has an edge relationship with "coughing up phlegm", "dry cough" and "hemoptysis" respectively, indicating that "cough" has a subordinate relationship with "coughing up phlegm", "dry cough" and "hemoptysis" respectively, and the edges between "coughing" and "coughing up phlegm", "dry cough" and "hemoptysis" respectively point to "coughing up phlegm", "dry cough" and "hemoptysis", indicating that "coughing up phlegm", "dry cough" and "hemoptysis" are sub-symptoms of "cough".

[0082] Similarly, "coughing up phlegm" and "blood in phlegm" have an edge relationship, indicating that "coughing up phlegm" and "blood in phlegm" have a subordinate relationship. The edge between "coughing up phlegm" and "blood in phlegm" points to "blood in phlegm", indicating that "blood in phlegm" is a sub-symptom of "coughing up phlegm".

[0083] Similarly, "hemoptysis" and "blood in sputum" have an edge relationship, indicating that "hemoptysis" and "blood in sputum" have a subordinate relationship. The edge between "hemoptysis" and "blood in sputum" points to "blood in sputum", indicating that "blood in sputum" is also a sub-symptom of "hemoptysis".

[0084] After determining the symptom relationship spectrum corresponding to each preset symptom, the fusion probability of each preset symptom occurring in each preset disease can be updated based on the symptom relationship spectrum to obtain the target probability of each preset symptom occurring in each preset disease. For example, the target probability of preset disease k occurring in preset symptom j can be where f z Represents the update function, which is used to fusion probability of preset disease k with preset symptom j update process.

[0085] Specifically, for each preset disease and each preset symptom, in response to the fusion probability of the preset symptom occurring in the preset disease being greater than the first threshold and the symptom relationship spectrum not including the child node of the third node (the node representing the preset symptom in the symptom relationship spectrum), the fusion probability of the preset symptom occurring in the preset disease is determined as the target probability of the preset symptom occurring in the preset disease.

[0086] That is, the fusion probability of the preset symptom j appearing in the preset disease k When the probability of the preset disease k appearing with the preset symptom j is greater than the first threshold ∈ and the symptom relationship spectrum does not include the sub-symptom of the preset symptom j, the fusion probability of the preset disease k appearing with the preset symptom j is Determine the target probability D of the preset symptom j for the preset disease k kj , which can be specifically expressed by the following formula:

[0087]

[0088] Among them, ∈ is the first threshold, child(j) represents the sub-symptom of the preset symptom j, that is, the child node of the node representing the preset symptom j, and child(j) is empty, which means that the symptom relationship spectrum does not include the node representing the sub-symptom of the preset symptom j.

[0089] Specifically, for each preset disease and each preset symptom, in response to the fusion probability of the preset symptom occurring in the preset disease being greater than a first threshold value, and the symptom relationship spectrum including a child node of the third node, the maximum fusion probability of the fusion probability of the preset symptom occurring in the preset disease and the target probabilities of the preset symptoms represented by each child node of the third node occurring in the preset disease (i.e., the sub-symptoms of the preset symptom) is determined as the target probability of the preset symptom occurring in the preset disease. Specifically, the maximum value of the fusion probability of the preset symptom occurring in the preset disease and the target probability of each sub-symptom of the preset symptom occurring in the preset disease can be determined as the target probability of the preset symptom occurring in the preset disease.

[0090] That is, the fusion probability of the preset symptom j appearing in the preset disease k When the probability of the preset disease k appearing with the preset symptom j is greater than the first threshold ∈ and the symptom relationship spectrum does not include the sub-symptom of the preset symptom j, the fusion probability of the preset disease k appearing with the preset symptom j is The maximum value of the target probability of each sub-symptom of the preset symptom occurring in the preset disease is determined as the target probability D of the preset disease k occurring in the preset symptom j. kj , which can be specifically expressed by the following formula:

[0091]

[0092] Among them, ∈ is the first threshold, and child(j) represents the sub-symptom set of the preset symptom j, that is, the child node of the node representing the preset symptom j. represents the target probability of sub-symptom j' of the preset symptom j occurring in the preset disease k, It represents the fusion probability of sub-symptom j' of preset symptom j occurring in preset disease k. If child(j) is not empty, it means that the symptom relationship spectrum includes nodes representing sub-symptoms of preset symptom j.

[0093] Specifically, for each preset disease and each preset symptom, in response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node in the symptom relationship spectrum including at least one fifth node, the target probability of the preset symptom occurring in the preset disease is determined based on the target probability of the preset symptoms represented by each sub-node of the third node (i.e., each sub-symptom of the preset symptom).

[0094] Among them, the fourth node includes each child node of the third node, and nodes indirectly associated with the third node (such as all descendant nodes of the third node), and the fusion probability of the preset symptoms represented by each fifth node of the preset disease is greater than the first threshold.

[0095] The above process can be specifically expressed by the following formula:

[0096]

[0097] Where ∈ is the first threshold, and Des(j) represents the sub-symptoms of symptom j and the symptoms indirectly associated with symptom j, i.e., the set of all descendant nodes of the node representing symptom j (all child nodes of the node representing symptom j and nodes associated with the node representing symptom j through at least one node). In this case, j' represents the fifth node, i.e., the node among all descendant nodes of the node representing symptom j for which the fusion probability of the corresponding symptom occurring in the disease is greater than the first threshold.

[0098] Among them, f p(j, k) represents the process of determining the target probability of the preset disease k appearing with the preset symptom j based on the target probability of the preset symptoms represented by each child node of the third node (i.e., each sub-symptom of the preset symptom j) appearing with the preset disease j, and

[0099]

[0100] Among them, based on the target probability of the preset symptoms represented by each sub-node of the third node of the preset disease j (that is, the sub-symptoms of the preset symptom j), the process of determining the target probability of the preset disease k appearing with the preset symptom j can also be determined based on the noisy-or function, which will not be repeated here.

[0101] Specifically, for each preset disease and each preset symptom, in response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the symptom relationship spectrum does not include the child node of the third node (that is, it does not include the sub-symptom of the preset symptom), or the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node does not include the fifth node (among the preset symptoms represented by all descendant nodes of the node representing the preset symptom, there is no preset symptom with a corresponding fusion probability greater than the first threshold), the preset probability is determined as the target probability of the preset symptom occurring in the preset disease.

[0102] The above process can be specifically expressed by the following formula:

[0103]

[0104] Wherein, P(j) is a preset probability, which can be a preset value close to 0 and is not limited here.

[0105] Based on this, the target probability of the preset disease k presenting the preset symptom j can be determined by the following formula:

[0106]

[0107] Optionally, for any symptom (such as cough), if the patient has a sub-symptom of the symptom (coughing up phlegm), the patient will inevitably have the symptom (cough). Based on this, after determining the symptom relationship spectrum corresponding to each preset symptom, the superordinate symptoms corresponding to some of the sub-symptoms can be filled in, and for each preset disease, the fusion probability of the superordinate symptom occurring in the preset disease can be assigned, and the value is greater than or equal to the fusion probability of any sub-symptom of the superordinate symptom occurring in the disease. In this case, the superordinate sub-symptom can also be determined as a preset symptom, and the fusion probability of each preset symptom occurring in each preset disease can be further updated based on the filled symptom relationship spectrum to obtain the target probability of each preset symptom occurring in each preset disease. The specific updating method will not be repeated here.

[0108] Step S24: Determine at least one target symptom of the patient to be predicted, and determine the target disease suffered by the patient to be predicted based on each target symptom and the target disease characteristic spectrum.

[0109] In some feasible embodiments, when determining the target disease suffered by the patient to be predicted, the target matching degree between the patient to be predicted and the above-mentioned various preset diseases can be determined based on at least one target symptom and target disease characteristic spectrum of the patient to be predicted.

[0110] Specifically, after determining the target matching degrees between the patient to be predicted and various preset diseases, the preset diseases corresponding to one or more target matching degrees with the largest values may be determined as the target diseases of the patient to be predicted.

[0111] Alternatively, for each preset disease, a combination score corresponding to the preset disease can be determined, specifically including the target matching degree between the patient to be predicted and the preset disease and the ranking score of the preset disease among all preset diseases. For example, for preset disease t, the combination score corresponding to preset disease t is:

[0112] p t =(K d -r t ,s t ),

[0113] Among them, s t Indicates the target matching degree between the patient to be predicted and the preset disease, r t ∈{1,2,3,…,K d} is the score ranking of the preset disease t among all preset diseases. When it ranks first, r t =1.

[0114] In response to K d -r t Greater than the ranking threshold, s t If the matching degree is greater than the threshold, it is determined that the patient to be predicted suffers from the preset disease t.

[0115] In some feasible embodiments, when determining the target match between the patient to be predicted and each predetermined disease, for each predetermined disease, a disease-discriminating weight for each predetermined symptom may be determined based on the target disease profile. The greater the disease-discriminating weight of any predetermined symptom, the stronger the disease-discriminating ability of that predetermined symptom.

[0116] The weight for distinguishing each preset symptom from the disease can be determined based on the principle of inverse document frequency (IDF). Specifically, for each preset symptom, each preset disease can be regarded as a document, and each preset symptom can be regarded as a word. When the target probability of the preset symptom occurring in the preset disease is greater than a preset probability threshold, it is determined that the preset disease includes (the document) the preset symptom (word). In this case, the weight for distinguishing the preset symptom from the disease can be determined based on the total number of preset diseases and the total number of preset diseases that include the preset symptom.

[0117] The disease discrimination weight IDF(j) of the preset symptom j can be determined in the following way:

[0118]

[0119] Among them, N c Indicates the number of preset diseases, represents an exemplary function, α is the probability threshold, D kj represents the target probability of the preset disease k presenting the preset symptom j, D kj Indicates that B is greater than the probability threshold kj The value is 1, otherwise the value is 0. kj When the value is 1, it means that the preset disease k includes the preset disease j, and when the value is 0, it means that the preset disease k does not include the preset disease j. It can represent the total number of preset diseases including the preset symptom j.

[0120] For each preset disease, after determining the distinguishing weight of each preset symptom for the disease based on the target disease characteristic spectrum, the target matching degree between the patient to be predicted and the preset disease can be determined based on the distinguishing weight of each target symptom corresponding to each preset symptom.

[0121] Specifically, the first symptom in each target symptom, wherein each target symptom does not include any sub-symptom of the first symptom. That is, the most fine-grained (lowest-level) target symptom in each target symptom is determined as the first symptom, and there is no situation in which one first symptom in each first symptom is a sub-symptom of another first symptom.

[0122] Furthermore, the first preset symptom among the preset symptoms can be determined, and the target probability of each first preset symptom of the preset disease is greater than the above probability threshold. That is, for the preset disease k, the first preset symptom among the preset symptoms can be expressed as: F k ={j|B kj =1,j=1,2,3,…N b}, N b is the total number of preset symptoms.

[0123] Furthermore, the intersection of each first symptom and each first preset symptom can be determined, and based on the discrimination weight corresponding to each symptom in the intersection of the first symptoms, the target matching degree between the patient to be predicted and the preset disease can be determined.

[0124] For a preset disease k, the target matching degree between the patient to be predicted and the preset disease k can be determined based on the following formula:

[0125]

[0126] Where Q represents the target symptom set, f del (Q) represents the process of determining the first symptoms of each item, f del (Q)∩F k represents the first symptom intersection, and g is the symptom in the first symptom intersection.

[0127] Alternatively, in order to further improve the matching relationship between the patient to be predicted and each predicted disease, for each predicted disease, the first matching degree between the patient to be predicted and the predicted disease can be determined based on the discrimination weights corresponding to the symptoms in the first symptom intersection. That is, for the preset disease k, S(Q→F k ) is determined as the first matching degree between the patient to be predicted and the predicted disease k. The second preset symptoms in each first preset symptom can be further determined, wherein each first preset symptom does not include any sub-symptom of the first preset symptom. That is, the finest-grained (lowest) first preset symptom in each first preset symptom is determined as the second preset symptom, and there is no situation in each second preset symptom that a second preset symptom is a sub-symptom of another second preset symptom. Among them, the second preset symptom set can be expressed as f del (F k ).

[0128] Furthermore, the intersection of each second preset symptom and the second symptom of each target symptom can be determined, and based on the discrimination weight corresponding to each symptom in the second symptom intersection, the second matching degree between the patient to be predicted and the preset disease can be determined. For the preset disease k, the second matching degree between the patient to be predicted and the preset disease k can be determined based on the following formula:

[0129]

[0130] Where Q represents the target symptom set, f del (F k ) represents the process of determining each second preset symptom, f del (F k )∩Q represents the second symptom intersection, and x is the symptom in the second symptom intersection.

[0131] For each preset disease, after determining the target matching degree between the patient to be predicted and the preset disease, the target matching degree between the patient to be predicted and the preset disease can be determined based on the corresponding first matching degree and second matching degree.

[0132] For example, the average of the first matching degree and the second matching degree can be determined as the target matching degree, which is not limited here. For the preset disease k, the target matching degree between the patient to be predicted and the preset disease k is It can be determined based on the following formula:

[0133]

[0134] In some feasible embodiments, when determining the target matching degree between the patient to be predicted and each preset disease, for each preset disease, the second probability that any patient suffers from the preset disease can be determined first, and then based on the target probability of each preset symptom of the preset disease, the target matching degree between the patient to be predicted and the preset disease can be determined.

[0135] The second probability that any patient suffers from the preset disease is a priori probability, which can be specifically determined based on existing statistical data on the prevalence of various diseases.

[0136] Specifically, the first symptom in each target symptom can be determined first, specifically the first symptom in each target symptom, wherein each target symptom does not include any sub-symptom of the first symptom. That is, the finest-grained (lowest) target symptom in each target symptom is determined as the first symptom, and there is no situation in each first symptom that a first symptom is a sub-symptom of another first symptom. If the target symptom set is Q, then the first symptom set is Q', and Q'=f del (Q), f del (Q) indicates the process of determining each first symptom.

[0137] Furthermore, in the case where the preset symptoms include the first symptoms, the third probability of the preset disease simultaneously appearing with the first symptoms may be determined based on the target probability of the preset disease appearing with each preset symptom.

[0138] The target probability of each first symptom of the preset disease appearing can be determined based on the characteristic spectrum of the target disease, and then the third probability of the preset disease appearing at the same time can be determined based on the target probability of each first symptom of the preset disease, and the third probability can be determined as the target matching degree between the patient to be predicted and the preset disease.

[0139] For a predetermined disease k, if the target probabilities of each first symptom of the predetermined disease k occurring are independent of each other, the third probability of the predetermined disease k simultaneously occurring with each first symptom can be determined based on the following method:

[0140]

[0141] Where m is the number of first symptoms, The first symptom j appears for the preset disease k m The target probability of the first symptom set is Q′.

[0142] In this embodiment of the present application, Π represents a continuous multiplication operation.

[0143] Among them, the target matching degree between the patient to be predicted and the preset disease can be expressed as S(Q,k)=P(y=k|Q′), which is used to express the probability that the patient to be predicted suffers from disease k when the first symptoms appear.

[0144] In some feasible implementations, when determining the target matching degree between the patient to be predicted and various preset diseases, the target symptoms of the patient to be predicted can also be input into the disease prediction model to obtain the target matching degree between the patient to be predicted and various preset diseases.

[0145] Among them, the above disease prediction model is obtained by training based on the characteristic spectrum of the target disease.

[0146] Specifically, when training a disease prediction model, a first training sample set can be constructed based on each disease information set. The first training sample set includes multiple first training samples. Each first training sample includes symptoms that occur when a patient suffers from a disease. Each first training sample is annotated with a sample label. The sample label of each first training sample represents the actual disease suffered by the corresponding patient.

[0147] Furthermore, based on the characteristic spectrum of the target disease, a second training sample corresponding to each first training sample can be constructed. Each second training sample also includes symptoms of a patient with a history of a disease, and each second training sample is also annotated with a sample label. The sample label of each second training sample represents the actual disease suffered by the corresponding patient.

[0148] Each group of first training samples and second training samples corresponds to the same patient and disease, but to different symptoms.

[0149] For each first training sample, a target probability of each symptom in the disease corresponding to the first training sample occurring can be determined based on the target disease profile. Furthermore, symptoms in the first training sample with a target probability less than a construction threshold are determined as the corresponding symptoms in the second training sample, and the patient and disease corresponding to the first training sample are determined as the patient and disease corresponding to the second training sample. Thus, while keeping the patient and disease corresponding to the first training sample unchanged, the corresponding second training sample is constructed by changing the symptoms in the first training sample.

[0150] For each first training sample, since the target probabilities of each symptom in the disease corresponding to the first training sample are independent of each other, a construction threshold can be determined for each symptom. For each symptom, if the target probability corresponding to the symptom is less than the corresponding construction threshold, the symptom is retained; otherwise, the symptom is removed, thereby obtaining the corresponding symptoms in the second training sample. This further improves the model's training effectiveness by ensuring that each symptom in the second training sample has a lower probability of occurring in the corresponding disease.

[0151] The construction threshold corresponding to each symptom may be the same or different, and there is no restriction here.

[0152] After each second training sample is determined, if the symptom in the second training sample is a sub-symptom of other superordinate symptoms, the superordinate symptom corresponding to the symptom is also determined as a symptom in the second training sample.

[0153] Furthermore, after constructing each second training sample, each symptom in each first training sample and each second training sample is input into the initial model to obtain the predicted matching degree between the patient corresponding to each first training sample and each preset disease, and the predicted disease of the patient corresponding to each first training sample is determined based on the predicted matching degree between the patient corresponding to each first training sample and each preset disease, and the predicted disease of the patient corresponding to each second training sample is determined based on the predicted matching degree between the patient corresponding to each second training sample and each preset disease, and the training loss value is determined based on the actual disease represented by the sample label of each first training sample and each second training sample, and the predicted disease corresponding to each first training sample and each second training sample, and the initial model is trained based on the training loss value and each first training sample and each second training sample until the training loss value meets the training end condition and the training is stopped, and the model at the time of stopping the training is determined as the disease prediction model.

[0154] During the training process, for each second training sample, the sample label of the second training sample is updated every preset time based on the predicted disease of the second training sample at the corresponding moment.

[0155] The above training loss value can be determined based on the following method:

[0156]

[0157] Among them, y i and y i 'represent the real disease corresponding to the first training sample and the real disease corresponding to the second training sample, respectively, N represents the predicted disease of the first training sample and the second training sample.v and N v ′ represents the number of samples of the first training sample and the second training sample respectively. is the cross entropy loss between the true disease and the predicted disease corresponding to the first training sample, is the cross entropy loss between the true disease and the predicted disease corresponding to the second training sample.

[0158] Here, t represents the preset time interval, and β(t) is a parameter that increases with training. At the beginning of training, the initial model is not accurate enough, and β(t) is small. In the later stages of training, the initial model is able to assign the correct label to the second training sample, and β(t) is large. The initial model uses pseudo samples (second training samples) to further improve its generalization ability.

[0159] Based on this model training approach, semi-supervised training of disease prediction models can be achieved with relatively few training samples, improving both training effectiveness and prediction accuracy. Furthermore, for diseases with low incidence or rare diseases, a rich set of training samples can be constructed and trained to produce a disease prediction model, effectively achieving accurate predictions for these diseases.

[0160] Among them, the various disease information sets and various training samples in the embodiments of the present application can be stored in a database, cloud storage or blockchain, and the specific storage can be determined based on the actual application scenario requirements and is not limited here.

[0161] In short, the database can be thought of as an electronic filing cabinet—a place for storing electronic files. In this application, it can be used to store various disease information sets and various training samples. Blockchain is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a string of data blocks generated through cryptographic methods. In this application, each data block in the blockchain can store various disease information sets and various training samples. Cloud storage is a new concept that extends and develops from the concept of cloud computing. It refers to the use of cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (also called storage nodes) in the network through application software or application interfaces to work together to store various disease information sets and various training samples.

[0162] The data processing and computing processes involved in the embodiments of this application can be performed based on cloud computing and other methods. Among them, cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing. Based on cloud computing, the efficiency of data processing and computing in the embodiments of this application can be improved.

[0163] The training method for the disease prediction model provided in the embodiments of this application can be implemented based on machine learning (ML) in artificial intelligence (AI). Machine learning (ML) specifically studies how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance, thereby implementing the training method for the disease prediction model based on semi-supervised learning in the embodiments of this application.

[0164] Among them, artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0165] The disease prediction method and the prediction method of the disease prediction model provided in the embodiments of the present application can be executed by any terminal device or server. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Among them, the terminal device can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this.

[0166] See also Figure 8 , Figure 8 Schematic diagram of the structure of the disease prediction device provided in the embodiment of the present application. The disease prediction device provided in the embodiment of the present application includes:

[0167] An information determination module 81 is configured to determine at least one disease information set, each of which includes at least one disease suffered by a patient and the symptoms that occur;

[0168] a probability determination module 82 for determining, for each of the disease information sets, an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set;

[0169] a probability fusion module 83 for determining a target disease characteristic spectrum based on each of the initial disease characteristic spectra, wherein the target disease characteristic spectrum represents a target probability of each preset symptom occurring in each preset disease, each target probability being determined based on at least one first probability of the preset disease corresponding to the target probability occurring in each of the initial disease characteristic spectra, each of the preset diseases being one of the diseases corresponding to each of the disease information sets, and each of the preset symptoms being one of the symptoms corresponding to each of the disease information sets;

[0170] The disease prediction module 84 is used to determine at least one target symptom of the patient to be predicted, and determine the target disease suffered by the patient to be predicted based on the target symptoms and the target disease characteristic spectrum.

[0171] In some feasible implementations, for each of the disease information sets, the probability determination module 82 is configured to:

[0172] For each disease and each symptom in the disease information set, determining a first number of patients suffering from the disease and a second number of patients suffering from the disease and experiencing the symptom in the disease information set, and determining a first probability that the disease experiences the symptom based on the first number and the second number;

[0173] Based on the first probabilities corresponding to each disease in the disease information set, an initial disease characteristic spectrum corresponding to the disease information set is determined.

[0174] In some feasible implementations, the probability fusion module 83 is configured to:

[0175] For each of the preset diseases and each of the preset symptoms, determining a first disease characteristic spectrum in each of the initial disease characteristic spectrums that includes a first probability corresponding to the preset disease, and determining a fused probability that the preset disease will appear with the preset symptom based on the third number of each of the first disease characteristic spectrums and the first probability that the preset disease will appear with the preset symptom in each of the first disease characteristic spectrums;

[0176] Based on the fusion probability of each of the above-mentioned preset symptoms occurring in each of the above-mentioned preset diseases, a target disease characteristic spectrum is determined.

[0177] In some feasible implementations, the probability fusion module 83 is configured to:

[0178] Determine a symptom relationship spectrum corresponding to each of the preset symptoms, wherein a node in the symptom relationship spectrum represents one of the preset symptoms, and a preset symptom represented by a subnode of a node in the symptom relationship spectrum is a subsymptom of the preset symptom represented by the node;

[0179] Based on the symptom relationship spectrum, the fusion probability of each of the preset symptoms occurring in each of the preset diseases is updated to obtain the target probability of each of the preset symptoms occurring in each of the preset diseases.

[0180] In some feasible implementations, for each of the above-mentioned preset diseases and each of the above-mentioned preset symptoms, the above-mentioned probability fusion module 83 is used to:

[0181] In response to the fused probability of the preset symptom occurring in the preset disease being greater than a first threshold and the symptom relationship spectrum not including a child node of the third node, determining the fused probability of the preset symptom occurring in the preset disease as a target probability of the preset symptom occurring in the preset disease, where the third node is a node representing the preset symptom in the symptom relationship spectrum;

[0182] In response to the fused probability of the preset symptom occurring in the preset disease being greater than the first threshold, and the symptom relationship spectrum including a child node of the third node, determining the maximum fused probability of the fused probability of the preset symptom occurring in the preset disease and the target probabilities of the preset symptoms occurring in the preset disease represented by each child node of the third node as the target probability of the preset symptom occurring in the preset disease;

[0183] In response to the fused probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node in the symptom relationship spectrum including at least one fifth node, determining a target probability of the preset symptom occurring in the preset disease based on the target probability of the preset symptom represented by each child node of the third node occurring in the preset disease, wherein the fourth node includes each child node of the third node and a node indirectly associated with the third node, and the fused probability of the preset symptom occurring in the preset disease represented by each fifth node is greater than the first threshold;

[0184] In response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the above-mentioned first threshold, and the above-mentioned symptom relationship spectrum does not include the child node of the above-mentioned third node, or the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the above-mentioned first threshold, and the fourth node in the above-mentioned symptom relationship spectrum does not include the above-mentioned fifth node, the preset probability is determined as the target probability of the preset symptom occurring in the preset disease.

[0185] In some feasible implementations, the disease prediction module 84 is configured to:

[0186] Based on each of the above target symptoms and the above target disease characteristic spectrum, determining the target matching degree between the above patient to be predicted and each of the above preset diseases;

[0187] Based on the target matching degree between the patient to be predicted and the various preset diseases, the target disease suffered by the patient to be predicted is determined.

[0188] In some feasible implementations, for each of the above-mentioned preset diseases, the above-mentioned disease prediction module 84 is used to:

[0189] Based on the target disease characteristic spectrum, determining the discrimination weight of each of the above-mentioned preset symptoms for the disease, and based on the discrimination weight corresponding to each of the above-mentioned preset symptoms and each of the above-mentioned target symptoms, determining the target matching degree between the above-mentioned patient to be predicted and the preset disease;

[0190] Determining a second probability that any patient suffers from the predetermined disease, and determining a target matching degree between the patient to be predicted and the predetermined disease based on the second probability and a target probability of each predetermined symptom of the predetermined disease occurring;

[0191] Each of the above target symptoms is input into a disease prediction model to obtain a target matching degree between the above patient to be predicted and the preset disease, wherein the above disease prediction model is trained based on the above target disease characteristic spectrum.

[0192] In some feasible implementations, for each of the above-mentioned preset diseases, the above-mentioned disease prediction module 84 is used to:

[0193] Determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom;

[0194] Determine the first symptom intersection of each of the above-mentioned first symptoms and the first preset symptom in each of the above-mentioned preset symptoms, and determine the target matching degree between the above-mentioned patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the above-mentioned first symptom intersection, wherein the target probability of the occurrence of each of the above-mentioned first preset symptoms in the preset disease is greater than the probability threshold.

[0195] In some feasible implementations, for each of the above-mentioned preset diseases, the above-mentioned disease prediction module 84 is used to:

[0196] Determining a first matching degree between the patient to be predicted and the preset disease based on the discrimination weights corresponding to the symptoms in the intersection of the first symptoms;

[0197] Determining a second preset symptom in each of the first preset symptoms, wherein each of the first preset symptoms does not include a sub-symptom of the second preset symptom;

[0198] Determining the intersection of the second preset symptoms and the second symptoms of the target symptoms, and determining a second matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the intersection of the second symptoms;

[0199] Based on the first matching degree and the second matching degree, a target matching degree between the patient to be predicted and the preset disease is determined.

[0200] In some feasible implementations, for each of the above-mentioned preset diseases, the above-mentioned disease prediction module 84 is used to:

[0201] Determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom;

[0202] Determining a third probability of the simultaneous occurrence of each of the first symptoms of the predetermined disease based on the target probability of the occurrence of each of the predetermined symptoms of the predetermined disease;

[0203] Based on the second probability and the third probability, a target matching degree between the patient to be predicted and the preset disease is determined.

[0204] In some feasible embodiments, the disease prediction model is obtained through training using a training device, and the training device is used to:

[0205] Constructing a plurality of first training samples based on each of the disease information sets, each of the first training samples including symptoms of a patient suffering from a disease, each of the first training samples being annotated with a sample label, and each of the sample labels of the first training samples representing the actual disease suffered by the corresponding patient;

[0206] Based on the target disease feature spectrum, construct a second training sample corresponding to each of the first training samples, each of the second training samples including symptoms of a patient suffering from a disease, each of the second training samples being annotated with a sample label, and each sample label of the second training set representing the actual disease suffered by the corresponding patient;

[0207] The first training sample and the second training sample in each group correspond to the same patient and disease, but correspond to different symptoms;

[0208] Input each symptom in each of the above-mentioned first training samples and each of the above-mentioned second training samples into the initial model, obtain the predicted matching degree between the patient corresponding to each of the above-mentioned first training samples and each of the above-mentioned second training samples and each of the above-mentioned preset diseases, determine the predicted disease of the patient corresponding to each of the above-mentioned first training samples based on the predicted matching degree between the patient corresponding to each of the above-mentioned first training samples and each of the above-mentioned preset diseases, determine the predicted disease of the patient corresponding to each of the above-mentioned second training samples based on the predicted matching degree between the patient corresponding to each of the above-mentioned second training samples and each of the above-mentioned preset diseases, determine the training loss value based on the actual disease represented by the sample label of each of the above-mentioned first training samples and each of the above-mentioned second training samples, and the predicted disease corresponding to each of the above-mentioned first training samples and each of the above-mentioned second training samples, train the above-mentioned initial model based on the training loss value, each of the above-mentioned first training samples and each of the above-mentioned second training samples, and stop the training when the above-mentioned training loss value meets the training end condition, and determine the model at the time of stopping the training as the above-mentioned disease prediction model;

[0209] During the training process, for each of the second training samples, the sample label of the second training sample is updated every preset time based on the predicted disease of the second training sample at the corresponding moment.

[0210] In some feasible implementations, for each of the first training samples, the training device is configured to:

[0211] Based on the target disease characteristic spectrum, determining a target probability of each symptom in the first training sample occurring in the disease corresponding to the first training sample;

[0212] Based on the symptoms in the first training sample whose target probability is less than the construction threshold, the patients and diseases corresponding to the first training sample, a second training sample corresponding to the first training sample is constructed.

[0213] In a specific implementation, the above device can execute the above-mentioned functions through its built-in functional modules. Figure 2 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.

[0214] See also Figure 9 , Figure 9 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 9As shown, the electronic device 900 in this embodiment may include: a processor 901, a network interface 904 and a memory 905. In addition, the above-mentioned electronic device 900 may also include: an object interface 903, and at least one communication bus 902. The communication bus 902 is used to realize the connection and communication between these components. The object interface 903 may include a display screen (Display), a keyboard (Keyboard), and the object interface 903 may optionally include a standard wired interface and a wireless interface. The network interface 904 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 905 may be a high-speed RAM memory or a non-volatile memory (NVM), such as at least one disk memory. The memory 905 may optionally also be at least one storage device located away from the aforementioned processor 901. As Figure 9 As shown, the memory 905 as a computer-readable storage medium may include an operating system, a network communication module, an object interface module, and a device control application program.

[0215] exist Figure 9 In the electronic device 900 shown, the network interface 904 can provide network communication functions; the object interface 903 is mainly used to provide an input interface for the object; and the processor 901 can be used to call the device control application stored in the memory 905 to achieve:

[0216] Determine at least one disease information set, each of which includes at least one disease suffered by a patient and the symptoms that occur;

[0217] For each of the above disease information sets, determining an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set;

[0218] Determine a target disease characteristic spectrum based on each of the initial disease characteristic spectra, wherein the target disease characteristic spectrum represents a target probability of each preset symptom occurring in each preset disease, each of the target probabilities being determined based on at least one first probability of the preset disease corresponding to the target probability occurring in each of the initial disease characteristic spectra, each of the preset diseases being one of the diseases corresponding to each of the disease information sets, and each of the preset symptoms being one of the symptoms corresponding to each of the disease information sets;

[0219] Determine at least one target symptom of the patient to be predicted, and based on the target symptoms and the target disease characteristic spectrum, determine the target disease suffered by the patient to be predicted.

[0220] In some feasible implementations, for each of the disease information sets, the processor 901 is configured to:

[0221] For each disease and each symptom in the disease information set, determining a first number of patients suffering from the disease and a second number of patients suffering from the disease and experiencing the symptom in the disease information set, and determining a first probability that the disease experiences the symptom based on the first number and the second number;

[0222] Based on the first probabilities corresponding to each disease in the disease information set, an initial disease characteristic spectrum corresponding to the disease information set is determined.

[0223] In some feasible implementations, the processor 901 is configured to:

[0224] For each of the preset diseases and each of the preset symptoms, determining a first disease characteristic spectrum in each of the initial disease characteristic spectrums that includes a first probability corresponding to the preset disease, and determining a fused probability that the preset disease will appear with the preset symptom based on the third number of each of the first disease characteristic spectrums and the first probability that the preset disease will appear with the preset symptom in each of the first disease characteristic spectrums;

[0225] Based on the fusion probability of each of the above-mentioned preset symptoms occurring in each of the above-mentioned preset diseases, a target disease characteristic spectrum is determined.

[0226] In some feasible implementations, the processor 901 is configured to:

[0227] Determine a symptom relationship spectrum corresponding to each of the preset symptoms, wherein a node in the symptom relationship spectrum represents one of the preset symptoms, and a preset symptom represented by a subnode of a node in the symptom relationship spectrum is a subsymptom of the preset symptom represented by the node;

[0228] Based on the symptom relationship spectrum, the fusion probability of each of the preset symptoms occurring in each of the preset diseases is updated to obtain the target probability of each of the preset symptoms occurring in each of the preset diseases.

[0229] In some feasible implementations, for each of the above-mentioned preset diseases and each of the above-mentioned preset symptoms, the processor 901 is configured to:

[0230] In response to the fused probability of the preset symptom occurring in the preset disease being greater than a first threshold and the symptom relationship spectrum not including a child node of the third node, determining the fused probability of the preset symptom occurring in the preset disease as a target probability of the preset symptom occurring in the preset disease, where the third node is a node representing the preset symptom in the symptom relationship spectrum;

[0231] In response to the fused probability of the preset symptom occurring in the preset disease being greater than the first threshold, and the symptom relationship spectrum including a child node of the third node, determining the maximum fused probability of the fused probability of the preset symptom occurring in the preset disease and the target probabilities of the preset symptoms occurring in the preset disease represented by each child node of the third node as the target probability of the preset symptom occurring in the preset disease;

[0232] In response to the fused probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node in the symptom relationship spectrum including at least one fifth node, determining a target probability of the preset symptom occurring in the preset disease based on the target probability of the preset symptom represented by each child node of the third node occurring in the preset disease, wherein the fourth node includes each child node of the third node and a node indirectly associated with the third node, and the fused probability of the preset symptom occurring in the preset disease represented by each fifth node is greater than the first threshold;

[0233] In response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the above-mentioned first threshold, and the above-mentioned symptom relationship spectrum does not include the child node of the above-mentioned third node, or the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the above-mentioned first threshold, and the fourth node in the above-mentioned symptom relationship spectrum does not include the above-mentioned fifth node, the preset probability is determined as the target probability of the preset symptom occurring in the preset disease.

[0234] In some feasible implementations, the processor 901 is configured to:

[0235] Based on each of the above target symptoms and the above target disease characteristic spectrum, determining the target matching degree between the above patient to be predicted and each of the above preset diseases;

[0236] Based on the target matching degree between the patient to be predicted and the various preset diseases, the target disease suffered by the patient to be predicted is determined.

[0237] In some feasible implementations, for each of the above-mentioned preset diseases, the processor 901 is configured to:

[0238] Based on the target disease characteristic spectrum, determining the discrimination weight of each of the above-mentioned preset symptoms for the disease, and based on the discrimination weight corresponding to each of the above-mentioned preset symptoms and each of the above-mentioned target symptoms, determining the target matching degree between the above-mentioned patient to be predicted and the preset disease;

[0239] Determining a second probability that any patient suffers from the predetermined disease, and determining a target matching degree between the patient to be predicted and the predetermined disease based on the second probability and a target probability of each predetermined symptom of the predetermined disease occurring;

[0240] Each of the above target symptoms is input into a disease prediction model to obtain a target matching degree between the above patient to be predicted and the preset disease, wherein the above disease prediction model is trained based on the above target disease characteristic spectrum.

[0241] In some feasible implementations, for each of the above-mentioned preset diseases, the processor 901 is configured to:

[0242] Determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom;

[0243] Determine the first symptom intersection of each of the above-mentioned first symptoms and the first preset symptom in each of the above-mentioned preset symptoms, and determine the target matching degree between the above-mentioned patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the above-mentioned first symptom intersection, wherein the target probability of the occurrence of each of the above-mentioned first preset symptoms in the preset disease is greater than the probability threshold.

[0244] In some feasible implementations, for each of the above-mentioned preset diseases, the processor 901 is configured to:

[0245] Determining a first matching degree between the patient to be predicted and the preset disease based on the discrimination weights corresponding to the symptoms in the intersection of the first symptoms;

[0246] Determining a second preset symptom in each of the first preset symptoms, wherein each of the first preset symptoms does not include a sub-symptom of the second preset symptom;

[0247] Determining the intersection of the second preset symptoms and the second symptoms of the target symptoms, and determining a second matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the intersection of the second symptoms;

[0248] Based on the first matching degree and the second matching degree, a target matching degree between the patient to be predicted and the preset disease is determined.

[0249] In some feasible implementations, the processor 901 is configured to:

[0250] Determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom;

[0251] Determining a third probability of the simultaneous occurrence of each of the first symptoms of the predetermined disease based on the target probability of the occurrence of each of the predetermined symptoms of the predetermined disease;

[0252] Based on the second probability and the third probability, a target matching degree between the patient to be predicted and the preset disease is determined.

[0253] In some feasible implementations, the processor 901 is configured to:

[0254] Constructing a plurality of first training samples based on each of the disease information sets, each of the first training samples including symptoms of a patient suffering from a disease, each of the first training samples being annotated with a sample label, and each of the sample labels of the first training samples representing the actual disease suffered by the corresponding patient;

[0255] Based on the target disease feature spectrum, construct a second training sample corresponding to each of the first training samples, each of the second training samples including symptoms of a patient suffering from a disease, each of the second training samples being annotated with a sample label, and each sample label of the second training set representing the actual disease suffered by the corresponding patient;

[0256] The first training sample and the second training sample in each group correspond to the same patient and disease, but correspond to different symptoms;

[0257] Input each symptom in each of the above-mentioned first training samples and each of the above-mentioned second training samples into the initial model, obtain the predicted matching degree between the patient corresponding to each of the above-mentioned first training samples and each of the above-mentioned second training samples and each of the above-mentioned preset diseases, determine the predicted disease of the patient corresponding to each of the above-mentioned first training samples based on the predicted matching degree between the patient corresponding to each of the above-mentioned first training samples and each of the above-mentioned preset diseases, determine the predicted disease of the patient corresponding to each of the above-mentioned second training samples based on the predicted matching degree between the patient corresponding to each of the above-mentioned second training samples and each of the above-mentioned preset diseases, determine the training loss value based on the actual disease represented by the sample label of each of the above-mentioned first training samples and each of the above-mentioned second training samples, and the predicted disease corresponding to each of the above-mentioned first training samples and each of the above-mentioned second training samples, train the above-mentioned initial model based on the training loss value, each of the above-mentioned first training samples and each of the above-mentioned second training samples, and stop the training when the above-mentioned training loss value meets the training end condition, and determine the model at the time of stopping the training as the above-mentioned disease prediction model;

[0258] During the training process, for each of the second training samples, the sample label of the second training sample is updated every preset time based on the predicted disease of the second training sample at the corresponding moment.

[0259] In some feasible implementations, for each of the first training samples, the processor 901 is configured to:

[0260] Based on the target disease characteristic spectrum, determining a target probability of each symptom in the first training sample occurring in the disease corresponding to the first training sample;

[0261] Based on the symptoms in the first training sample whose target probability is less than the construction threshold, the patients and diseases corresponding to the first training sample, a second training sample corresponding to the first training sample is constructed.

[0262] It should be understood that in some feasible embodiments, the processor 901 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store device type information.

[0263] In a specific implementation, the electronic device 900 can execute the above-mentioned functions through its built-in functional modules. Figure 2 For the implementation methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.

[0264] The present invention also provides a computer-readable storage medium that stores a computer program and is executed by a processor to implement Figure 2 For the methods provided in each step, please refer to the implementation methods provided in the above steps for details, which will not be repeated here.

[0265] The above-mentioned computer-readable storage medium can be an internal storage unit of the device or electronic device provided in any of the aforementioned embodiments, such as a hard disk or memory of an electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above-mentioned computer-readable storage medium can also include a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc. Further, the computer-readable storage medium can also include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0266] The present invention provides a computer program product, which includes a computer program, which is executed by a processor. Figure 2 The methods provided in each step.

[0267] The terms "first," "second," and the like in the claims, specification, and drawings of this application are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or electronic device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or electronic device. Reference herein to an "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The presence of such a phrase in various locations in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive with other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments. The term "and / or," as used in this specification and the appended claims, refers to any and all possible combinations of one or more of the associated listed items, including, but not limited to, those combinations.

[0268] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the above description generally describes the components and steps of each example according to their functions. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0269] The above disclosure is only a preferred embodiment of the present application and cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A disease prediction method, characterized in that: The method comprises: Determine at least one disease information set, each of the disease information sets including at least one disease suffered by a patient and symptoms; For each of the disease information sets, determining an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set; For each preset disease and each preset symptom, determining a first disease characteristic spectrum including a first probability corresponding to the preset disease in each of the initial disease characteristic spectrums, and determining a fusion probability of the preset disease appearing with the preset symptom based on the third number of each of the first disease characteristic spectrums and the first probability of the preset disease appearing with the preset symptom in each of the first disease characteristic spectrums; each of the preset diseases is a disease among the diseases corresponding to each of the disease information sets, and each of the preset symptoms is a symptom among the symptoms corresponding to each of the disease information sets; Determine a symptom relationship spectrum corresponding to each of the preset symptoms, wherein a node in the symptom relationship spectrum represents a preset symptom, and a preset symptom represented by a subnode of a node in the symptom relationship spectrum is a subsymptom of the preset symptom represented by the node; Based on the symptom relationship spectrum, the fusion probability of each preset symptom occurring in each preset disease is updated to obtain a target disease characteristic spectrum, wherein the target disease characteristic spectrum represents the target probability of each preset symptom occurring in each preset disease; At least one target symptom of the patient to be predicted is determined, and based on each of the target symptoms and the target disease characteristic spectrum, the target disease suffered by the patient to be predicted is determined.

2. The method according to claim 1, characterized in that For each of the disease information sets, determining the initial disease characteristic spectrum corresponding to the disease information set includes: For each disease and each symptom in the disease information set, determining a first number of patients suffering from the disease and a second number of patients suffering from the disease and experiencing the symptom in the disease information set, and determining a first probability that the disease experiences the symptom based on the first number and the second number; Based on the first probabilities corresponding to each disease in the disease information set, an initial disease characteristic spectrum corresponding to the disease information set is determined.

3. The method according to claim 1, characterized in that For each of the preset diseases and each of the preset symptoms, the fusion probability of the preset disease occurring with the preset symptom is updated based on the symptom relationship spectrum to obtain the target probability of the preset disease occurring with the preset symptom, including: In response to the fused probability of the preset symptom occurring in the preset disease being greater than a first threshold and the symptom relationship spectrum not including a child node of a third node, determining the fused probability of the preset symptom occurring in the preset disease as a target probability of the preset symptom occurring in the preset disease, wherein the third node is a node representing the preset symptom in the symptom relationship spectrum; In response to the fused probability of the preset symptom occurring in the preset disease being greater than the first threshold, and the symptom relationship spectrum including a child node of the third node, determining the maximum fused probability of the fused probability of the preset symptom occurring in the preset disease and the target probabilities of the preset symptoms represented by each child node of the third node occurring in the preset disease as the target probability of the preset symptom occurring in the preset disease; In response to the fused probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node in the symptom relationship spectrum including at least one fifth node, determining a target probability of the preset symptom occurring in the preset disease based on the target probability of the preset symptom represented by each child node of the third node occurring in the preset disease, wherein the fourth node includes each child node of the third node and a node indirectly associated with the third node, and the fused probability of the preset symptom occurring in the preset disease represented by each fifth node is greater than the first threshold; In response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold and the symptom relationship spectrum not including the child node of the third node, or the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold and the fourth node in the symptom relationship spectrum not including the fifth node, the preset probability is determined as the target probability of the preset symptom occurring in the preset disease.

4. The method according to claim 1, wherein Determining the target disease suffered by the patient to be predicted based on each of the target symptoms and the target disease characteristic spectrum includes: Determining the target matching degree between the patient to be predicted and each of the preset diseases based on each of the target symptoms and the target disease characteristic spectrum; Based on the target matching degree between the patient to be predicted and various preset diseases, the target disease suffered by the patient to be predicted is determined.

5. The method according to claim 4, characterized in that For each of the preset diseases, based on the target symptoms and the target disease profile, determining the target matching degree between the patient to be predicted and the preset disease includes at least one of the following: Based on the target disease characteristic spectrum, determining the distinguishing weight of each of the preset symptoms for the disease, and determining the target matching degree between the patient to be predicted and the preset disease based on the distinguishing weight corresponding to each of the preset symptoms and each of the target symptoms; Determining a second probability that any patient suffers from the predetermined disease, and determining a target matching degree between the patient to be predicted and the predetermined disease based on the second probability and a target probability of each predetermined symptom occurring in the predetermined disease; Each target symptom is input into a disease prediction model to obtain a target matching degree between the patient to be predicted and the preset disease, wherein the disease prediction model is trained based on the characteristic spectrum of the target disease.

6. The method according to claim 5, characterized in that For each of the preset diseases, determining the target matching degree between the patient to be predicted and the preset disease based on the discrimination weights corresponding to each of the preset symptoms and each of the target symptoms includes: determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom; Determine the first symptom intersection of each of the first symptoms and the first preset symptom in each of the preset symptoms, and determine the target matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection, wherein the target probability of the occurrence of each of the first preset symptoms in the preset disease is greater than a probability threshold.

7. The method according to claim 6, characterized in that For each of the preset diseases, determining the target matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection includes: Determining a first matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection; determining a second preset symptom among each of the first preset symptoms, wherein each of the first preset symptoms does not include a sub-symptom of the second preset symptom; Determining a second symptom intersection between each of the second preset symptoms and each of the target symptoms, and determining a second matching degree between the patient to be predicted and the preset disease based on a discrimination weight corresponding to each symptom in the second symptom intersection; Based on the first matching degree and the second matching degree, a target matching degree between the patient to be predicted and the preset disease is determined.

8. The method according to claim 5, characterized in that Determining a target matching degree between the patient to be predicted and the preset disease based on the second probability and the target probability of each preset symptom occurring in the preset disease includes: determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom; Determining, based on the target probability of each of the predetermined symptoms occurring in the predetermined disease, a third probability of each of the first symptoms occurring simultaneously in the predetermined disease; Based on the second probability and the third probability, a target matching degree between the patient to be predicted and the preset disease is determined.

9. The method according to claim 5, characterized in that The disease prediction model is trained based on the following method: constructing a plurality of first training samples based on each of the disease information sets, each of the first training samples including symptoms of a patient suffering from a disease, each of the first training samples being annotated with a sample label, and the sample label of each of the first training samples representing the actual disease suffered by the corresponding patient; Based on the target disease feature spectrum, constructing a second training sample corresponding to each of the first training samples, each of the second training samples includes symptoms that occur when a patient suffers from a disease, each of the second training samples is annotated with a sample label, and each sample label of the second training set represents the actual disease suffered by the corresponding patient; The first training sample and the second training sample in each group correspond to the same patient and disease, but correspond to different symptoms; Inputting each symptom in each of the first training samples and each of the second training samples into an initial model, obtaining a predicted matching degree between the patient corresponding to each of the first training samples and each of the second training samples and each of the preset diseases, determining a predicted disease for the patient corresponding to each of the first training samples based on the predicted matching degree between the patient corresponding to each of the first training samples and each of the preset diseases, determining a predicted disease for the patient corresponding to each of the second training samples based on the predicted matching degree between the patient corresponding to each of the second training samples and each of the preset diseases, determining a training loss value based on the actual disease represented by the sample label of each of the first training samples and each of the second training samples, and the predicted disease corresponding to each of the first training samples and each of the second training samples, training the initial model based on the training loss value, each of the first training samples, and each of the second training samples until the training is stopped when the training loss value meets a training end condition, and determining the model at the time of stopping the training as the disease prediction model; During the training process, for each of the second training samples, the sample label of the second training sample is updated every preset time based on the predicted disease of the second training sample at the corresponding moment.

10. The method according to claim 9, characterized in that For each of the first training samples, constructing a second training sample corresponding to the first training sample based on the target disease feature spectrum includes: Determining, based on the target disease characteristic spectrum, a target probability of each symptom in the first training sample occurring in the disease corresponding to the first training sample; Based on the symptoms in the first training sample whose target probability is less than the construction threshold, the patients and diseases corresponding to the first training sample, a second training sample corresponding to the first training sample is constructed.

11. A disease prediction device, characterized in that: The device comprises: An information determination module, configured to determine at least one disease information set, each of which includes at least one disease suffered by a patient and the symptoms that occur; a probability determination module, configured to determine, for each of the disease information sets, an initial disease characteristic spectrum corresponding to the disease information set, wherein the initial disease characteristic spectrum represents a first probability of each symptom in the disease information set occurring for each disease in the disease information set; a probability fusion module, configured to determine, for each preset disease and each preset symptom, a first disease characteristic spectrum including a first probability corresponding to the preset disease in each of the initial disease characteristic spectra, and determine a fusion probability of the preset disease appearing with the preset symptom based on the third number of each of the first disease characteristic spectra and the first probability of the preset disease appearing with the preset symptom in each of the first disease characteristic spectra; each of the preset diseases being one of the diseases corresponding to each of the disease information sets, and each of the preset symptoms being one of the symptoms corresponding to each of the disease information sets; The probability fusion module is used to determine the symptom relationship spectrum corresponding to each of the preset symptoms, wherein a node in the symptom relationship spectrum represents a preset symptom, and the preset symptom represented by the subnode of a node in the symptom relationship spectrum is a sub-symptom of the preset symptom represented by the node; The probability fusion module is configured to update the fusion probability of each of the preset symptoms occurring in each of the preset diseases based on the symptom relationship spectrum to obtain a target disease characteristic spectrum, wherein the target disease characteristic spectrum represents the target probability of each of the preset symptoms occurring in each of the preset diseases; The disease prediction module is used to determine at least one target symptom of the patient to be predicted, and determine the target disease suffered by the patient to be predicted based on each of the target symptoms and the target disease characteristic spectrum.

12. The device according to claim 11, characterized in that For each of the disease information sets, when the probability determination module determines the initial disease characteristic spectrum corresponding to the disease information set, it is used to: For each disease and each symptom in the disease information set, determining a first number of patients suffering from the disease and a second number of patients suffering from the disease and experiencing the symptom in the disease information set, and determining a first probability that the disease experiences the symptom based on the first number and the second number; Based on the first probabilities corresponding to each disease in the disease information set, an initial disease characteristic spectrum corresponding to the disease information set is determined.

13. The device according to claim 11, characterized in that For each of the preset diseases and each of the preset symptoms, the probability fusion module updates the fusion probability of the preset disease appearing with the preset symptom based on the symptom relationship spectrum, and when obtaining the target probability of the preset disease appearing with the preset symptom, is used to: In response to the fused probability of the preset symptom occurring in the preset disease being greater than a first threshold and the symptom relationship spectrum not including a child node of a third node, determining the fused probability of the preset symptom occurring in the preset disease as a target probability of the preset symptom occurring in the preset disease, wherein the third node is a node representing the preset symptom in the symptom relationship spectrum; In response to the fused probability of the preset symptom occurring in the preset disease being greater than the first threshold, and the symptom relationship spectrum including a child node of the third node, determining the maximum fused probability of the fused probability of the preset symptom occurring in the preset disease and the target probabilities of the preset symptoms represented by each child node of the third node occurring in the preset disease as the target probability of the preset symptom occurring in the preset disease; In response to the fused probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold, and the fourth node in the symptom relationship spectrum including at least one fifth node, determining a target probability of the preset symptom occurring in the preset disease based on the target probability of the preset symptom represented by each child node of the third node occurring in the preset disease, wherein the fourth node includes each child node of the third node and a node indirectly associated with the third node, and the fused probability of the preset symptom occurring in the preset disease represented by each fifth node is greater than the first threshold; In response to the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold and the symptom relationship spectrum not including the child node of the third node, or the fusion probability of the preset symptom occurring in the preset disease being less than or equal to the first threshold and the fourth node in the symptom relationship spectrum not including the fifth node, the preset probability is determined as the target probability of the preset symptom occurring in the preset disease.

14. The device according to claim 11, characterized in that When the disease prediction module determines the target disease suffered by the patient to be predicted based on each of the target symptoms and the target disease characteristic spectrum, it is used to: Determining the target matching degree between the patient to be predicted and each of the preset diseases based on each of the target symptoms and the target disease characteristic spectrum; Based on the target matching degree between the patient to be predicted and various preset diseases, the target disease suffered by the patient to be predicted is determined.

15. The device according to claim 14, characterized in that For each of the preset diseases, the disease prediction module determines the target matching degree between the patient to be predicted and the preset disease based on each of the target symptoms and the target disease characteristic spectrum, and is used to: Based on the target disease characteristic spectrum, determining the distinguishing weight of each of the preset symptoms for the disease, and determining the target matching degree between the patient to be predicted and the preset disease based on the distinguishing weight corresponding to each of the preset symptoms and each of the target symptoms; Determining a second probability that any patient suffers from the predetermined disease, and determining a target matching degree between the patient to be predicted and the predetermined disease based on the second probability and a target probability of each predetermined symptom occurring in the predetermined disease; Each target symptom is input into a disease prediction model to obtain a target matching degree between the patient to be predicted and the preset disease, wherein the disease prediction model is trained based on the characteristic spectrum of the target disease.

16. The device according to claim 15, characterized in that For each of the preset diseases, the disease prediction module determines the target matching degree between the patient to be predicted and the preset disease based on the discrimination weights corresponding to the preset symptoms and the target symptoms, and is used to: determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom; Determine the first symptom intersection of each of the first symptoms and the first preset symptom in each of the preset symptoms, and determine the target matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection, wherein the target probability of the occurrence of each of the first preset symptoms in the preset disease is greater than a probability threshold.

17. The device according to claim 16, characterized in that For each of the preset diseases, the disease prediction module determines the target matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection, and is used to: Determining a first matching degree between the patient to be predicted and the preset disease based on the discrimination weight corresponding to each symptom in the first symptom intersection; determining a second preset symptom among each of the first preset symptoms, wherein each of the first preset symptoms does not include a sub-symptom of the second preset symptom; Determining a second symptom intersection between each of the second preset symptoms and each of the target symptoms, and determining a second matching degree between the patient to be predicted and the preset disease based on a discrimination weight corresponding to each symptom in the second symptom intersection; Based on the first matching degree and the second matching degree, a target matching degree between the patient to be predicted and the preset disease is determined.

18. The device according to claim 15, characterized in that The disease prediction module is configured to determine the target matching degree between the patient to be predicted and the preset disease based on the second probability and the target probability of each preset symptom of the preset disease: determining a first symptom among each of the target symptoms, wherein each of the target symptoms does not include a sub-symptom of the first symptom; Determining, based on the target probability of each of the predetermined symptoms occurring in the predetermined disease, a third probability of each of the first symptoms occurring simultaneously in the predetermined disease; Based on the second probability and the third probability, a target matching degree between the patient to be predicted and the preset disease is determined.

19. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor and the memory are connected to each other; The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 10 when calling the computer program.

20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 10.

21. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Disease prediction device and equipment, and symptom information processing method, device and equipment

    CN112768064A