Disease risk analysis method, device, electronic device and computer storage medium
By classifying the training data set and aggregating weak classifiers, a high-precision disease model is established, and data matching is accelerated through index relationships, the accuracy and efficiency of disease risk analysis in the prior art is solved, and more efficient and accurate disease risk analysis is achieved.
Patent Information
- Application Number
- CN202010459737.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-05-26
AI Technical Summary
The prediction accuracy and inefficiency of disease risk analysis in the prior art makes it difficult to achieve high-precision and high-efficiency analysis.
By obtaining the training data set for classification, using the classification data set for training multiple weak classifiers, selecting the target weak classifier and aggregating it into a disease model. The user data set to be judged is preprocessed, an index relationship is established and matched, and a disease model is used for analysis and calculation to obtain disease analysis results.
The efficiency and accuracy of disease risk analysis are improved, the accuracy of the model is improved through the utilization of classified data sets, and the data matching and analysis process is accelerated through the establishment of index relationships.
Smart Images

Figure CN111696663B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a disease risk analysis method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the rise of big data, data processing technology has been applied to various fields. Today, when people are paying more and more attention to their physical health, there is no shortage of data processing technology in the medical field that uses various information data of patients to analyze people's health status and disease risks, thereby evaluating their health status.
[0003] In the prior art, the methods for analyzing the risk of disease mainly use highly interpretable methods such as logistic regression and decision trees. However, due to the large number of subjective factors in this method, the prediction accuracy is not high and the prediction efficiency is low. Therefore, how to achieve high-precision and high-efficiency analysis of the risk of disease has become an urgent problem to be solved. Summary of the invention
[0004] The present invention provides a disease risk analysis method, device, electronic device and computer-readable storage medium, the main purpose of which is to improve the efficiency and accuracy of disease risk analysis.
[0005] To achieve the above object, the present invention provides a method for analyzing disease risk, comprising:
[0006] Obtaining a training data set, and classifying the training data set to obtain a classified data set;
[0007] Using the classification data set to train a plurality of pre-constructed weak classifiers, and selecting a plurality of target weak classifiers from the trained plurality of weak classifiers, and aggregating the target weak classifiers into a disease model;
[0008] Acquire a user data set to be determined, and preprocess the user data set to be determined to obtain a target data set;
[0009] Establishing an index relationship between the target data set and the classification data set;
[0010] Matching the target data set with the classification data set according to the index relationship to obtain a matching data set;
[0011] The disease model is used to analyze and calculate the matching data set to obtain a disease analysis result.
[0012] Optionally, preprocessing the to-be-judged user data set to obtain a target data set includes:
[0013] Identify missing data in the to-be-determined user data set, and obtain a missing data set;
[0014] generating predicted data for missing data in the missing data set;
[0015] The predicted data is filled into the user data set to be judged to obtain the target data set.
[0016] Optionally, establishing an index relationship between the target data set and the classification data set includes:
[0017] Creating a category data table in a disease database according to the categories included in the classification data set;
[0018] Determine the target category to which the target data in the target data set belongs in the category data table;
[0019] An index relationship is established between the target data in the target data set and the classification data set according to the target category.
[0020] Optionally, matching the target data set with the classification data set according to the index relationship to obtain a matching data set includes:
[0021] Extracting characters from a plurality of target data in the target data set to generate a plurality of character data sets corresponding to the plurality of target data;
[0022] The multiple character data sets are matched with the classification data set through the index relationship to generate a matching data set.
[0023] Optionally, after obtaining the disease analysis result, the method further comprises:
[0024] Comparing the disease analysis result with a preset result threshold;
[0025] When the disease analysis result is less than or equal to the result threshold, sending a first treatment plan reminder;
[0026] When the disease analysis result is greater than the result threshold, a second treatment plan reminder is sent.
[0027] In order to solve the above problems, the present invention further provides a disease risk analysis device, the device comprising:
[0028] A data classification module is used to obtain a training data set and classify the training data set to obtain a classified data set;
[0029] A model training module, used to train a plurality of pre-constructed weak classifiers using the classification data set, and select a plurality of target weak classifiers from the trained plurality of weak classifiers, and aggregate the target weak classifiers into a disease model;
[0030] A target data acquisition module is used to acquire a user data set to be determined, and pre-process the user data set to be determined to obtain a target data set;
[0031] An index relationship establishment module, used to establish an index relationship between the target data set and the classification data set;
[0032] A data matching module, used to match the target data set with the classified data set according to the index relationship to obtain a matching data set;
[0033] The analysis and calculation module is used to use the disease model to analyze and calculate the matching data set to obtain a disease analysis result.
[0034] Optionally, the target data acquisition module is specifically used to:
[0035] Acquire a user data set to be determined, identify missing data in the user data set to be determined, and obtain a missing data set;
[0036] generating predicted data for missing data in the missing data set;
[0037] The predicted data is filled into the user data set to be judged to obtain the target data set.
[0038] Optionally, the index relationship establishing module is specifically used for:
[0039] Creating a category data table in the disease database according to the categories included in the classification data set;
[0040] Determine the target category to which the target data in the target data set belongs in the category data table;
[0041] An index relationship is established between the target data in the target data set and the classification data set according to the target category.
[0042] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0043] a memory storing at least one instruction; and
[0044] A processor executes instructions stored in the memory to implement any one of the disease risk analysis methods described above.
[0045] In order to solve the above problems, the present invention also provides a computer-readable storage medium, including a data storage area and a program storage area, the data storage area stores data created according to the use of blockchain nodes, and the program storage area stores a computer program; wherein, when the computer program is executed by a processor, an analysis method for disease risk as described in any one of the above items is implemented.
[0046] In the embodiment of the present invention, the acquired training data set is classified, and the classification data set obtained by classification is used to train multiple pre-constructed weak classifiers, and multiple target weak classifiers are selected from the multiple weak classifiers after training, and the weak classifiers are aggregated into a disease model; after obtaining the user data set to be judged, the data set to be judged is preprocessed to obtain a target data set, and an index relationship is established between the target data set and the classification data set; the target data set is matched with the classification data set according to the index relationship to obtain a matching data set; the matching data set is analyzed and calculated using the disease model to obtain a disease analysis result. By classifying the training data set before training the model, the efficiency of model training can be improved, and by training the base model with classification data sets of different categories, the accuracy of the disease model can be improved, which is conducive to improving the accuracy of disease model analysis; at the same time, when the user data set to be judged is analyzed, the target data set obtained by preprocessing the user data set to be judged is indexed with the classification data set, so as to quickly and accurately find the corresponding relationship between the target data set and the classification data set, determine the category corresponding to the target data set, and quickly identify the disease risk corresponding to the user data to be judged according to the category through the disease model. Therefore, the disease risk analysis method, device and computer-readable storage medium proposed in the present invention can achieve the purpose of improving the efficiency and accuracy of disease risk analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of a flow chart of a method for analyzing disease risk provided by an embodiment of the present invention;
[0048] Figure 2 A schematic diagram of a module of a disease risk analysis device provided by an embodiment of the present invention;
[0049] Figure 3 A schematic diagram of the internal structure of an electronic device for implementing a disease risk analysis method provided by an embodiment of the present invention;
[0050] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0051] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0052] The execution subject of the disease risk analysis method provided in the embodiment of the present application includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the method provided in the embodiment of the present application. In other words, the disease risk analysis method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc.
[0053] Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0054] The underlying blockchain platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining public and private key generation (account management), key management, and the maintenance of the correspondence between the user's real identity and the blockchain address (authority management), etc., and, under authorization, supervises and audits the transactions of certain real identities and provides risk control rule configuration (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and records valid requests to storage after consensus is reached. For a new business request, the basic service first adapts the interface for parsing and authentication (interface adaptation), and then encrypts the business information through the consensus algorithm (consensus management). The smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define the contract logic in a programming language and publish it to the blockchain (contract registration). According to the logic of the contract terms, the key or other events are called to trigger the execution and complete the contract logic. It also provides the function of contract upgrade and cancellation. The operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation and real-time status visualization output of the product during the product release process, such as alarm, network monitoring, node equipment health monitoring, etc.
[0055] The present invention provides a method for analyzing disease risk. Figure 1 FIG. 1 is a flow chart of a method for analyzing disease risk provided by an embodiment of the present invention. The method can be executed by a device, and the device can be implemented by software and / or hardware.
[0056] In this embodiment, the disease risk analysis method includes:
[0057] S1. Obtain a training data set, and classify the training data set to obtain a classified data set.
[0058] In the embodiment of the present invention, the training data set is a data set that records the user's disease history information, and the training data set includes but is not limited to: information related to physical condition (such as gender, age, allergy history, etc.), disease name, disease type, disease symptoms, and disease medication. The training data can be obtained from a database used by various hospitals to store patient data, and the database stores data after desensitization of patient data.
[0059] Furthermore, in an embodiment of the present invention, when the training data set is classified, the training data set is classified according to different features. For example, the training data in the training data set are classified according to the same disease type; or, the training data set is classified according to the same disease symptom, and the classified data in the classified data set corresponds to different diseases.
[0060] Preferably, in order to better utilize the classified data set later, the embodiment of the present invention stores the classified data set in a pre-built disease database, and the disease database may be a MySQL database, an Oracle database, or the like.
[0061] S2. Using the classification data set to train multiple pre-constructed weak classifiers, and selecting multiple target weak classifiers from the trained multiple weak classifiers, and aggregating the target weak classifiers into a disease model.
[0062] Specifically, in an optional embodiment of the present invention, the weak classifier is:
[0063] h(δ i ,p,θ)=pδ i <pθ
[0064] Among them, h(δ i ,p,θ) is the classification result of the weak classifier, δ i is the classified data in the classified data set, p is the indicator parameter of the preset inequality direction, and θ is the preset classification threshold.
[0065] In specific implementation, multiple different classification thresholds are preset for θ, thereby obtaining multiple weak classifiers. Multiple classification results can be obtained by classifying the classification data set using multiple weak classifiers, wherein each weak classifier corresponds to one classification result.
[0066] Optionally, in an embodiment of the present invention, after obtaining a plurality of weak classifiers and performing classification according to the weak classifiers, the plurality of pre-trained weak classifiers are selected using an error rate function according to the classification results to obtain a plurality of target weak classifiers.
[0067] The error rate function is:
[0068]
[0069] Among them, w i is the classification data set, y i is the classification result of the classification data in the classification data set.
[0070] Preferably, in the embodiment of the present invention, a plurality of weak classifiers having an error rate less than a preset error threshold are screened out as target weak classifiers.
[0071] Preferably, the number of the target weak classifiers is consistent with the number of categories of the classification data in the classification data set.
[0072] In detail, the disease model is as follows:
[0073]
[0074] Where t is the number of target weak classifiers, f k is the target weak classifier, F is the set of all target weak classifiers, is the output result of the disease model.
[0075] By training multiple weak classifiers and selecting the pre-trained multiple weak classifiers using an error rate function, a target weak classifier with higher accuracy can be obtained, and then the multiple weak classifiers with higher accuracy can be aggregated into a disease model to improve the accuracy of the disease model.
[0076] S3: Obtain a user data set to be determined, and preprocess the user data set to be determined to obtain a target data set.
[0077] In a preferred embodiment of the present invention, the user data set to be determined can be stored in a blockchain node.
[0078] Specifically, the present invention can use pre-edited Java statements to call the user data set to be determined from nodes used for one or more blockchains.
[0079] In an embodiment of the present invention, the user data set to be judged includes but is not limited to: information of the user to be judged (such as gender, age, etc.), historical diseases of the user to be judged, historical medication of the user to be judged, and historical disease symptoms of the user to be judged. The number of users to be judged can be one or more.
[0080] The preprocessing includes but is not limited to: data filling, data correction, data deletion, and data standardization.
[0081] Further, in an optional embodiment of the present invention, the preprocessing of the to-be-judged user data set to obtain the target data set includes:
[0082] Identify missing data in the to-be-determined user data set, and obtain a missing data set;
[0083] generating predicted data for missing data in the missing data set;
[0084] Fill the predicted data into the user data set to be judged to obtain the target data set. Preferably, the embodiment of the present invention can use pre-edited java statements to perform length detection on the user data to be judged in the user data set to be judged. The user data to be judged contains multiple attribute data with judged users and corresponding values. For example, the user data set to be judged contains age data of the user to be judged and the value corresponding to the age data; during the specific detection, the value corresponding to each attribute data in the user data to be judged is detected. When it is detected that the length of the value is not 0 or not null, it is determined that the value of the attribute data is not missing, and the detection continues; when it is detected that the length of the value is 0 or null, it is determined that the value of the attribute data is missing, and the attribute data and corresponding values of all missing values are obtained as a set of missing data, that is, a missing data set. Preferably, in the embodiment of the present invention, the generation of predicted data for missing data in the missing data set includes:
[0085] Use mice function to select the adjacent data of any missing data in the missing data set;
[0086] The mean of the adjacent data is calculated to obtain the predicted data.
[0087] In detail, the embodiment of the present invention uses the mice function, takes the position of any missing data in the missing data set in the user data set to be judged as the center point, sets a length threshold, selects adjacent data within the length threshold, and uses the following mean algorithm to calculate the mean of the adjacent data to obtain the predicted data Avg:
[0088]
[0089] Where V is the number of adjacent data, D v Any adjacent data.
[0090] The embodiment of the present invention can make the data to be judged more complete by filling in the missing user data to be judged, which is beneficial to improving the accuracy of model training.
[0091] S4. Establish an index relationship between the target data set and the classification data set.
[0092] Preferably, the step of establishing an index relationship between the target data set and the classification data set in the disease database comprises:
[0093] Creating a category data table in a disease database according to the categories included in the classification data set;
[0094] Determine the target category to which the target data in the target data set belongs in the category data table;
[0095] Establishing an index relationship between the target data in the target data set and the classification data set according to the target category. Preferably, establishing an index relationship between the target data in the target data set and the classification data set according to the target category means that searching the category data table based on any data such as the user's historical diseases, the user's historical medications, and the user's historical disease symptoms contained in the target data set, and classifying the target data in the target data set into corresponding categories according to the search results.
[0096] For example, when the classified data set is classified according to the disease symptoms contained in the classified data set, when performing the retrieval, the retrieval is performed according to the user's historical disease symptoms contained in the target data set, and the target data in the target data set and the corresponding category of the classified data in the classified data set are retrieved, which is the index relationship.
[0097] Furthermore, before establishing an index relationship between the target data set and the classification data set, the method described in the embodiment of the present invention also includes: transmitting the target data set to the disease database via the TCP / IP protocol, wherein the TCP / IP protocol is a data transmission protocol, and the data transmission interface of the disease database can be called according to the TCP / IP protocol, thereby facilitating the efficient transmission of the target data set to the disease database.
[0098] S5. Match the target data set with the classified data set according to the index relationship to obtain a matching data set.
[0099] Further, in another optional embodiment of the present invention, matching the target data set with the classified data set according to the index relationship to obtain a matching data set includes:
[0100] Extracting characters from a plurality of target data in the target data set to generate a plurality of character data sets corresponding to the plurality of target data;
[0101] The multiple character data sets are matched with the classification data set through the index relationship to generate a matching data set.
[0102] In detail, a preset character grabber may be used to extract characters from the target data in the target data set, wherein the character grabber is a python statement, and the python statement is used for character grabbing.
[0103] Specifically, preferably, after obtaining the multiple character data sets, the embodiment of the present invention matches the multiple character data sets with the classification data set through the index relationship to generate a matching data set, that is, finds the corresponding categories of the character data set and the classification data in the classification data set according to the index relationship. The matching data set includes the target data in the target data set and the classification data in the classification data set corresponding to the target data.
[0104] Preferably, the present invention further comprises performing array aggregation on the character data set using the following array aggregation algorithm to generate an array data set:
[0105]
[0106] Where J is the array data set, β i is a character in the character data set, and m is the number of characters in the character data set.
[0107] The character data set is array-aggregated to generate an array data set, and the data are aggregated together to further improve the efficiency of subsequent data processing.
[0108] S6. Analyze and calculate the matching data set using the disease model to obtain a disease analysis result.
[0109] In the embodiment of the present invention, the analysis result is the probability that the user to be determined corresponding to the target data in the target data set suffers from the disease corresponding to the classification data in the classification data set.
[0110] Preferably, the embodiment of the present invention uses the following analysis algorithm to perform the analysis calculation to obtain the analysis result
[0111]
[0112] Among them, x i is the matching data in the matching data set, t is the number of weak classifiers in the disease model, f t (xi ) is the output of the weak classifier.
[0113] Furthermore, the embodiment of the present invention also includes sending a treatment plan reminder based on the disease analysis result.
[0114] In detail, after obtaining the disease analysis result, the method further includes: comparing the disease analysis result with a preset result threshold;
[0115] When the disease analysis result is less than or equal to the result threshold, sending a first treatment plan reminder;
[0116] When the disease analysis result is greater than the result threshold, a second treatment plan reminder is sent.
[0117] When sending a treatment plan reminder, the reminder may be directly sent to the user to be determined corresponding to the user data set to be determined.
[0118] In this embodiment, the first treatment plan and the second treatment plan may be different treatment plans corresponding to different disease severity levels.
[0119] Furthermore, the treatment plan reminder includes an analysis of the cause of the disease.
[0120] This embodiment sends treatment plan reminders through disease analysis results, which helps relevant personnel quickly obtain personalized demand information.
[0121] In the embodiment of the present invention, the acquired training data set is classified, and the classification data set obtained by classification is used to train multiple pre-constructed weak classifiers, and multiple target weak classifiers are selected from the multiple weak classifiers after training, and the weak classifiers are aggregated into a disease model; after obtaining the user data set to be judged, the data set to be judged is preprocessed to obtain a target data set, and an index relationship is established between the target data set and the classification data set; the target data set is matched with the classification data set according to the index relationship to obtain a matching data set; the matching data set is analyzed and calculated using the disease model to obtain a disease analysis result. By classifying the training data set before training the model, the efficiency of model training can be improved, and by training the base model with classification data sets of different categories, the accuracy of the disease model can be improved, which is conducive to improving the accuracy of disease model analysis; at the same time, when the user data set to be judged is analyzed, the target data set obtained by preprocessing the user data set to be judged is indexed with the classification data set, so as to quickly and accurately find the corresponding relationship between the target data set and the classification data set, determine the category corresponding to the target data set, and quickly identify the disease risk corresponding to the user data to be judged according to the category through the disease model.
[0122] like Figure 2Shown is a module schematic diagram of the disease risk analysis device of the present invention.
[0123] The disease risk analysis device 100 of the present invention can be installed in an electronic device. According to the functions to be implemented, the disease risk analysis device can include a data classification module 101, a model training module 102, a target data acquisition module 103, an index relationship establishment module 104, a data matching module 105 and an analysis and calculation module 106. The module described in the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, which are stored in the memory of the electronic device.
[0124] In this embodiment, the functions of each module / unit are as follows:
[0125] The data classification module 101 is used to obtain a training data set and classify the training data set to obtain a classified data set;
[0126] The model training module 102 is used to train a plurality of pre-constructed weak classifiers using the classification data set, and select a plurality of target weak classifiers from the trained plurality of weak classifiers, and aggregate the target weak classifiers into a disease model;
[0127] The target data acquisition module 103 is used to acquire a user data set to be determined, and pre-process the user data set to be determined to obtain a target data set;
[0128] The index module 104 is used to establish an index relationship between the target data set and the classification data set;
[0129] The data matching module 105 is used to match the target data set with the classified data set according to the index relationship to obtain a matching data set;
[0130] The analysis and calculation module 106 is used to use the disease model to perform analysis and calculation on the matching data set to obtain a disease analysis result.
[0131] In detail, the specific implementation of each module of the disease risk analysis device is as follows:
[0132] The data classification module 101 is used to obtain a training data set and classify the training data set to obtain a classified data set.
[0133] In the embodiment of the present invention, the training data set is a data set that records the user's disease history information, and the training data set includes but is not limited to: information related to physical condition (such as gender, age, allergy history, etc.), disease name, disease type, disease symptoms, and disease medication. The training data can be obtained from a database used by various hospitals to store patient data, and the database stores data after desensitization of patient data.
[0134] Furthermore, in an embodiment of the present invention, when the training data set is classified, the training data set is classified according to different features. For example, the training data in the training data set are classified according to the same disease type; or, the training data set is classified according to the same disease symptom, and the classified data in the classified data set corresponds to different diseases.
[0135] Preferably, in order to better utilize the classified data set later, the embodiment of the present invention stores the classified data set in a pre-built disease database, and the disease database may be a MySQL database, an Oracle database, or the like.
[0136] The model training module 102 is used to train multiple pre-constructed weak classifiers using the classification data set, and select multiple target weak classifiers from the trained multiple weak classifiers, and aggregate the target weak classifiers into a disease model.
[0137] Specifically, in an optional embodiment of the present invention, the weak classifier is:
[0138] h(δ i ,p,θ)=pδ i <pθ
[0139] Among them, h(δ i ,p,θ) is the classification result of the weak classifier, δ i is the classified data in the classified data set, p is the indicator parameter of the preset inequality direction, and θ is the preset classification threshold.
[0140] In specific implementation, multiple different classification thresholds are preset for θ, thereby obtaining multiple weak classifiers. Multiple classification results can be obtained by classifying the classification data set using multiple weak classifiers, wherein each weak classifier corresponds to one classification result.
[0141] Optionally, in an embodiment of the present invention, after obtaining a plurality of weak classifiers and performing classification according to the weak classifiers, the plurality of pre-trained weak classifiers are selected using an error rate function according to the classification results to obtain a plurality of target weak classifiers.
[0142] The error rate function is:
[0143]
[0144] Among them, w i is the classification data set, y i is the classification result of the classification data in the classification data set.
[0145] Preferably, in the embodiment of the present invention, a plurality of weak classifiers having an error rate less than a preset error threshold are screened out as target weak classifiers.
[0146] Preferably, the number of the target weak classifiers is consistent with the number of categories of the classification data in the classification data set.
[0147] In detail, the disease model is as follows:
[0148]
[0149] Where t is the number of target weak classifiers, f k is the target weak classifier, F is the set of all target weak classifiers, is the output result of the disease model.
[0150] By training multiple weak classifiers and selecting the pre-trained multiple weak classifiers using an error rate function, a target weak classifier with higher accuracy can be obtained, and then the multiple weak classifiers with higher accuracy can be aggregated into a disease model to improve the accuracy of the disease model.
[0151] The target data acquisition module 103 is used to acquire a user data set to be determined, and pre-process the user data set to be determined to obtain a target data set.
[0152] In a preferred embodiment of the present invention, the user data set to be determined can be stored in a blockchain node.
[0153] Specifically, the present invention can use pre-edited Java statements to call the user data set to be determined from nodes used for one or more blockchains.
[0154] In an embodiment of the present invention, the user data set to be judged includes but is not limited to: information of the user to be judged (such as gender, age, etc.), historical diseases of the user to be judged, historical medication of the user to be judged, and historical disease symptoms of the user to be judged. The number of users to be judged can be one or more.
[0155] The preprocessing includes but is not limited to: data filling, data correction, data deletion, and data standardization.
[0156] Furthermore, in an optional embodiment of the present invention, the target data acquisition module 103 is specifically used for:
[0157] Acquire a user data set to be determined, identify missing data in the user data set to be determined, and obtain a missing data set;
[0158] generating predicted data for missing data in the missing data set;
[0159] The predicted data is filled into the user data set to be judged to obtain the target data set.
[0160] Preferably, the embodiment of the present invention can use pre-edited java statements to perform length detection on the user data to be judged in the user data set to be judged, where the user data to be judged contains multiple attribute data with the user to be judged and corresponding values, such as the age data of the user to be judged and the value corresponding to the age data in the user data set to be judged; during the specific detection, the values corresponding to the various attribute data in the user data to be judged are detected, and when it is detected that the length of the value is not 0 or not null, it is determined that the value of the attribute data is not missing, and the detection continues; when it is detected that the length of the value is 0 or null, it is determined that the value of the attribute data is missing, and the attribute data and corresponding values of all missing values are obtained to obtain a set of missing data, that is, a missing data set.
[0161] Preferably, in an embodiment of the present invention, generating predicted data of missing data in the missing data set includes:
[0162] Use mice function to select the adjacent data of any missing data in the missing data set;
[0163] The mean of the adjacent data is calculated to obtain the predicted data.
[0164] In detail, the embodiment of the present invention uses the mice function, takes the position of any missing data in the missing data set in the user data set to be judged as the center point, sets a length threshold, selects adjacent data within the length threshold, and uses the following mean algorithm to calculate the mean of the adjacent data to obtain the predicted data Avg:
[0165]
[0166] Where V is the number of adjacent data, D v Any adjacent data.
[0167] The embodiment of the present invention can make the data to be judged more complete by filling in the missing user data to be judged, which is beneficial to improving the accuracy of model training.
[0168] The index relationship establishing module 104 is used to establish an index relationship between the target data set and the classification data set.
[0169] Preferably, the index relationship establishing module 104 is specifically used for:
[0170] Creating a category data table in a disease database according to the categories included in the classification data set;
[0171] Determine the target category to which the target data in the target data set belongs in the category data table;
[0172] An index relationship is established between the target data in the target data set and the classification data set according to the target category.
[0173] Preferably, establishing an index relationship between the target data in the target data set and the classification data set according to the target category means that according to any data such as the user's historical diseases, the user's historical medications, the user's historical disease symptoms contained in the target data set, searching the category data table, and classifying the target data in the target data set into corresponding categories according to the search results.
[0174] For example, when the classified data set is classified according to the disease symptoms contained in the classified data set, when performing the retrieval, the retrieval is performed according to the user's historical disease symptoms contained in the target data set, and the target data in the target data set and the corresponding category of the classified data in the classified data set are retrieved, which is the index relationship.
[0175] Furthermore, before establishing an index relationship between the target data set and the classification data set, the method described in the embodiment of the present invention also includes: transmitting the target data set to the disease database via the TCP / IP protocol, wherein the TCP / IP protocol is a data transmission protocol, and the data transmission interface of the disease database can be called according to the TCP / IP protocol, thereby facilitating the efficient transmission of the target data set to the disease database.
[0176] The data matching module 105 is used to match the target data set with the classified data set according to the index relationship to obtain a matching data set.
[0177] Further, in another optional embodiment of the present invention, the data matching module 105 is specifically used to: extract characters from multiple target data in the target data set to generate multiple character data sets corresponding to the multiple target data;
[0178] The multiple character data sets are matched with the classification data set through the index relationship to generate a matching data set.
[0179] In detail, a preset character grabber may be used to extract characters from the target data in the target data set, wherein the character grabber is a python statement, and the python statement is used for character grabbing.
[0180] Specifically, preferably, after obtaining the multiple character data sets, the embodiment of the present invention matches the multiple character data sets with the classification data set through the index relationship to generate a matching data set, that is, finds the corresponding categories of the character data set and the classification data in the classification data set according to the index relationship. The matching data set includes the target data in the target data set and the classification data in the classification data set corresponding to the target data.
[0181] Preferably, the present invention further comprises performing array aggregation on the character data set to generate an array data set using the following array aggregation algorithm:
[0182]
[0183] Where J is the array data set, β i is a character in the character data set, and m is the number of characters in the character data set.
[0184] The character data set is array-aggregated to generate an array data set, and the data are aggregated together to further improve the efficiency of subsequent data processing.
[0185] The analysis and calculation module 106 is used to use the disease model to perform analysis and calculation on the matching data set to obtain a disease analysis result.
[0186] In the embodiment of the present invention, the analysis result is the probability that the user to be determined corresponding to the target data in the target data set suffers from the disease corresponding to the classification data in the classification data set.
[0187] Preferably, the embodiment of the present invention uses the following analysis algorithm to perform the analysis calculation to obtain the analysis result
[0188]
[0189] Among them, x i is the matching data in the matching data set, t is the number of weak classifiers in the disease model, f t (x i ) is the output of the weak classifier.
[0190] Furthermore, the embodiment of the present invention also includes sending a treatment plan reminder based on the disease analysis result.
[0191] In detail, after obtaining the disease analysis result, the device further includes a message sending module, and the message sending module is used to:
[0192] Comparing the disease analysis result with a preset result threshold;
[0193] When the disease analysis result is less than or equal to the result threshold, sending a first treatment plan reminder;
[0194] When the disease analysis result is greater than the result threshold, a second treatment plan reminder is sent.
[0195] When sending a treatment plan reminder, the reminder may be directly sent to the user to be determined corresponding to the user data set to be determined.
[0196] In this embodiment, the first treatment plan and the second treatment plan may be different treatment plans corresponding to different disease severity levels.
[0197] Furthermore, the treatment plan reminder includes an analysis of the cause of the disease.
[0198] This embodiment sends treatment plan reminders through disease analysis results, which helps relevant personnel quickly obtain personalized demand information.
[0199] In the embodiment of the present invention, the acquired training data set is classified, and the classification data set obtained by classification is used to train multiple pre-constructed weak classifiers, and multiple target weak classifiers are selected from the multiple weak classifiers after training, and the weak classifiers are aggregated into a disease model; after obtaining the user data set to be judged, the data set to be judged is preprocessed to obtain a target data set, and an index relationship is established between the target data set and the classification data set; the target data set is matched with the classification data set according to the index relationship to obtain a matching data set; the matching data set is analyzed and calculated using the disease model to obtain a disease analysis result. By classifying the training data set before training the model, the efficiency of model training can be improved, and by training the base model with classification data sets of different categories, the accuracy of the disease model can be improved, which is conducive to improving the accuracy of disease model analysis; at the same time, when the user data set to be judged is analyzed, the target data set obtained by preprocessing the user data set to be judged is indexed with the classification data set, so as to quickly and accurately find the corresponding relationship between the target data set and the classification data set, determine the category corresponding to the target data set, and quickly identify the disease risk corresponding to the user data to be judged according to the category through the disease model.
[0200] like Figure 3 The figure is a schematic diagram of the structure of an electronic device for implementing the disease risk analysis method of the present invention.
[0201] The electronic device 1 may include a processor 10 , a memory 11 , and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10 , such as a disease risk analysis program 12 .
[0202] The memory 11 includes at least one type of readable storage medium, including flash memory, mobile hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. The memory 11 may be an internal storage unit of the electronic device 1 in some embodiments, such as a mobile hard disk of the electronic device 1. The memory 11 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 may not only be used to store application software and various types of data installed in the electronic device 1, such as the code of the disease risk analysis program 12, etc., but also be used to temporarily store data that has been output or is to be output.
[0203] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes or executes programs or modules stored in the memory 11 (for example, executing a disease risk analysis program, etc.), and calls data stored in the memory 11 to execute various functions of the electronic device 1 and process data.
[0204] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0205] In the embodiment of the present invention, the acquired training data set is classified, and the classification data set obtained by classification is used to train multiple pre-constructed weak classifiers, and multiple target weak classifiers are selected from the multiple weak classifiers after training, and the weak classifiers are aggregated into a disease model; after obtaining the user data set to be judged, the data set to be judged is preprocessed to obtain a target data set, and an index relationship is established between the target data set and the classification data set; the target data set is matched with the classification data set according to the index relationship to obtain a matching data set; the matching data set is analyzed and calculated using the disease model to obtain a disease analysis result. By classifying the training data set before training the model, the efficiency of model training can be improved, and by training the base model with classification data sets of different categories, the accuracy of the disease model can be improved, which is conducive to improving the accuracy of disease model analysis; at the same time, when the user data set to be judged is analyzed, the target data set obtained by preprocessing the user data set to be judged is indexed with the classification data set, so as to quickly and accurately find the corresponding relationship between the target data set and the classification data set, determine the category corresponding to the target data set, and quickly identify the disease risk corresponding to the user data to be judged according to the category through the disease model.
[0206] Figure 3 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0207] For example, although not shown, the electronic device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The electronic device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0208] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0209] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device 1 and to display a visual user interface.
[0210] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0211] The disease risk analysis program 12 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, and when running in the processor 10, it can achieve:
[0212] Obtaining a training data set, and classifying the training data set to obtain a classified data set;
[0213] Using the classification data set to train a plurality of pre-constructed weak classifiers, and selecting a plurality of target weak classifiers from the trained plurality of weak classifiers, and aggregating the target weak classifiers into a disease model;
[0214] Acquire a user data set to be determined, and preprocess the user data set to be determined to obtain a target data set;
[0215] Establishing an index relationship between the target data set and the classification data set;
[0216] Matching the target data set with the classification data set according to the index relationship to obtain a matching data set;
[0217] The disease model is used to analyze and calculate the matching data set to obtain a disease analysis result.
[0218] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0219] Furthermore, the computer-usable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0220] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0221] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0222] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0223] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0224] Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any associated diagram mark in the claims should not be regarded as limiting the claims involved.
[0225] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0226] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The second and other words are used to indicate names, but not to indicate any particular order.
[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for analyzing disease risk, characterized in that: The method comprises: Obtaining a training data set, and classifying the training data set to obtain a classified data set; Using the classification data set to train a plurality of pre-constructed weak classifiers to obtain a plurality of classification results, selecting a plurality of target weak classifiers from the plurality of trained weak classifiers according to the plurality of classification results and a preset error rate function, and aggregating the target weak classifiers into a disease model; Acquire a user data set to be determined, and preprocess the user data set to be determined to obtain a target data set; Establishing an index relationship between the target data set and the classification data set, including: creating a category data table in a disease database according to the categories included in the classification data set, determining the target category to which the target data in the target data set belongs in the category data table, and establishing the index relationship between the target data in the target data set and the classification data set according to the target category; Matching the target data set with the classification data set according to the index relationship to obtain a matching data set, including: extracting characters from a plurality of target data in the target data set to generate a plurality of character data sets corresponding to the plurality of target data, matching the plurality of character data sets with the classification data set through the index relationship to generate the matching data set; The disease model is used to analyze and calculate the matching data set to obtain a disease analysis result.
2. The method for analyzing disease risk according to claim 1, characterized in that: The preprocessing of the to-be-judged user data set to obtain a target data set includes: Identify missing data in the to-be-determined user data set, and obtain a missing data set; generating predicted data for missing data in the missing data set; The predicted data is filled into the user data set to be judged to obtain the target data set.
3. The method for analyzing disease risk according to any one of claims 1 to 2, characterized in that: After obtaining the disease analysis result, the method further comprises: Comparing the disease analysis result with a preset result threshold; When the disease analysis result is less than or equal to the result threshold, sending a first treatment plan reminder; When the disease analysis result is greater than the result threshold, a second treatment plan reminder is sent.
4. A disease risk analysis device, characterized in that: The device comprises: A data classification module is used to obtain a training data set and classify the training data set to obtain a classified data set; A model training module, used to train a plurality of pre-constructed weak classifiers using the classification data set to obtain a plurality of classification results, select a plurality of target weak classifiers from the plurality of trained weak classifiers according to the plurality of classification results and a preset error rate function, and aggregate the target weak classifiers into a disease model; A target data acquisition module is used to acquire a user data set to be determined, and pre-process the user data set to be determined to obtain a target data set; An index relationship establishment module, used to establish an index relationship between the target data set and the classification data set, including: creating a category data table in a disease database according to the categories included in the classification data set, determining the target category to which the target data in the target data set belongs in the category data table, and establishing the index relationship between the target data in the target data set and the classification data set according to the target category; A data matching module is used to match the target data set with the classification data set according to the index relationship to obtain a matching data set, including: extracting characters from a plurality of target data in the target data set to generate a plurality of character data sets corresponding to the plurality of target data, matching the plurality of character data sets with the classification data set through the index relationship to generate the matching data set; The analysis and calculation module is used to use the disease model to analyze and calculate the matching data set to obtain a disease analysis result.
5. The disease risk analysis device according to claim 4, characterized in that: The target data acquisition module is specifically used for: Acquire a user data set to be determined, identify missing data in the user data set to be determined, and obtain a missing data set; generating predicted data for missing data in the missing data set; The predicted data is filled into the user data set to be judged to obtain the target data set.
6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the disease risk analysis method according to any one of claims 1 to 3.
7. A computer-readable storage medium, comprising a data storage area and a program storage area, wherein the data storage area stores data created according to the use of a blockchain node, and the program storage area stores a computer program; wherein: When the computer program is executed by a processor, the method for analyzing disease risk according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Medical data classification method and device based on machine learning and computer equipment
CN110021439A
Intelligent heart disease detection method and device and computer readable storage medium
CN110363090A
Psychological scoring method and device, electronic equipment and storage medium
CN110910989A