Medical data risk identification method and device, electronic equipment and storage medium
By receiving underwriting query requests, obtaining the target user's medical data, and matching it with a decision tree model, the problem of low efficiency and low accuracy in identifying user medical risks by insurance companies is solved, achieving more efficient and accurate risk identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2026-03-24
AI Technical Summary
In existing technologies, insurance companies are inefficient and inaccurate in identifying users' medical risks, mainly due to the lack of professional medical risk control models and the inability to accurately obtain users' medical data, resulting in time-consuming, labor-intensive, and inaccurate manual identification.
By receiving underwriting query requests, the system identifies the target data source, retrieves medical data from the target data source, matches the initial underwriting decision tree model of the intended insurance product, configures the target underwriting decision tree model, uses the model for underwriting identification, and outputs the identification results, including generating incremental decision trees and comprehensive performance optimization decision trees.
It improves the efficiency and accuracy of medical risk identification by accurately acquiring medical data and optimizing decision tree models, thus achieving more efficient and accurate risk identification.
Smart Images

Figure CN114240677B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method, apparatus, electronic device, and storage medium for identifying risks in medical data. Background Technology
[0002] Currently, while people are paying attention to diet and exercise, they also purchase health insurance to mitigate future risks. As the number of health insurance buyers increases, how to effectively identify users' medical risks has become the biggest pain point for insurance companies.
[0003] In developing this invention, the inventors discovered that the identification of medical risks is primarily conducted manually by insurance staff, which is time-consuming, labor-intensive, and inefficient. Furthermore, most insurance staff lack medical backgrounds and cannot determine which diseases are insurable or not, leading to low accuracy in medical risk identification. Online underwriting also falls short: it lacks professional medical risk control models for analysis and judgment; it cannot accurately obtain users' medical data; and without medical data, it is impossible to analyze disease incidence rates, thus hindering the accurate identification of customers' medical risks. Summary of the Invention
[0004] In view of the above, it is necessary to propose a method, device, electronic device and storage medium for medical data risk identification, which can improve the efficiency and accuracy of risk identification.
[0005] A first aspect of the present invention provides a method for identifying risks in medical data, the method comprising:
[0006] Receive an underwriting query request for a target user, and determine the target data source based on the underwriting query request;
[0007] Obtain the target user's medical data from the target data source;
[0008] The target user's intended insurance products are obtained based on the underwriting query request, and an initial underwriting decision tree model corresponding to the type of the intended insurance product is matched.
[0009] The initial underwriting decision tree model is configured with rules to obtain the target underwriting decision tree model, which includes multiple decision trees;
[0010] The target underwriting decision tree model is used to perform underwriting identification based on the medical data, and the identification results of each decision tree in the target underwriting decision tree model are obtained. The underwriting query results are then output based on the identification results of multiple decision trees.
[0011] According to an optional embodiment of the present invention, determining the target data source based on the underwriting query request includes:
[0012] Obtain the target user's intended insurance region and authorized institution identifier from the underwriting query request;
[0013] The regional agency mapping table is used to determine the multiple data agency identifiers corresponding to the intended insurance region;
[0014] The target data organization identifier is obtained from the plurality of data organization identifiers based on the authorized organization identifier;
[0015] The data institution corresponding to the target data institution identifier is identified as the target data source.
[0016] According to an optional embodiment of the present invention, obtaining the target user's medical data from the target data source includes:
[0017] Obtain the initial medical dataset of the target user from the target data source. Each piece of initial medical data in the initial medical dataset includes patient description information and medical description information.
[0018] The initial medical dataset is sampled to obtain a medical sample set with the same data distribution as the initial medical dataset;
[0019] Medical description values are determined in the medical sample set such that the ratio of the number of initial medical data including the medical description values to the number of medical sample sets is greater than a first preset threshold.
[0020] Obtain patient description values corresponding to the medical description values from the medical sample set, such that the ratio of the number of initial medical data including the patient description values to the number of medical sample sets is greater than a second preset threshold.
[0021] Search the initial medical dataset for initial medical data that includes the medical description value but does not include the patient description value;
[0022] The initial medical data found will be used as the medical data of the target user.
[0023] According to an optional embodiment of the present invention, configuring rules on the initial underwriting decision tree model to obtain the target underwriting decision tree model includes:
[0024] Determine whether the target user's medical record identifier can be obtained from the medical data;
[0025] When the medical record identifier of the target user is obtained from the medical data, the number of the medical record identifiers is calculated;
[0026] Configure the number of decision trees according to the stated number;
[0027] Generate the target underwriting decision tree model based on the number of decision trees.
[0028] According to an optional embodiment of the present invention, the step of using the target underwriting decision tree model to perform underwriting identification based on the medical data, and obtaining the identification result of each decision tree in the target underwriting decision tree model, includes:
[0029] Obtain the diagnosis and treatment data corresponding to each medical record identifier in the medical data;
[0030] Multiple sets of diagnostic and treatment data are input into the target underwriting decision tree model for underwriting identification;
[0031] Obtain the identification result of each decision tree in the target underwriting decision tree model.
[0032] According to an optional embodiment of the present invention, after outputting the underwriting query result based on the recognition results of the multiple decision trees, the method further includes:
[0033] Retrieve incremental data within a predetermined time period;
[0034] Generate an incremental decision tree based on the incremental data;
[0035] Label prediction is performed on the incremental data based on the decision tree in the incremental decision tree and the decision tree in the target underwriting decision tree model to obtain the label prediction result;
[0036] The overall performance of the decision tree in the target underwriting decision tree model and the decision tree in the incremental decision tree is determined based on the label prediction results.
[0037] Based on the overall performance, a predetermined number of decision trees are selected from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree model as the updated target underwriting decision tree model.
[0038] According to an optional embodiment of the present invention, selecting a predetermined number of decision trees from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree includes:
[0039] The prediction accuracy of the decision tree for the incremental data is determined based on the results of the label prediction.
[0040] The time it takes to build the decision tree is used as the weight for the overall performance.
[0041] The prediction accuracy of the incremental data is sorted according to the weights.
[0042] The predetermined number of decision trees are obtained from the decision trees in the target underwriting decision tree model and the sorted prediction accuracy.
[0043] A second aspect of the present invention provides a medical data risk identification device, the device comprising:
[0044] The receiving module is used to receive underwriting query requests for target users and determine the target data source based on the underwriting query requests;
[0045] The acquisition module is used to acquire the target user's medical data from the target data source;
[0046] The matching module is used to obtain the target user's intended insurance products based on the underwriting query request, and match an initial underwriting decision tree model corresponding to the type of the intended insurance product;
[0047] The configuration module is used to configure rules for the initial underwriting decision tree model to obtain the target underwriting decision tree model, which includes multiple decision trees;
[0048] The identification module is used to perform underwriting identification based on the medical data using the target underwriting decision tree model, obtain the identification result of each decision tree in the target underwriting decision tree model, and output the underwriting query result based on the identification results of multiple decision trees.
[0049] A third aspect of the present invention provides an electronic device including a processor for implementing the medical data risk identification method when executing a computer program stored in a memory.
[0050] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the medical data risk identification method when executed by a processor.
[0051] In summary, the medical data risk identification method, device, electronic device, and storage medium of this invention determine the target data source through the underwriting query request of the target user, thereby accurately obtaining the target user's medical data from the target data source. Accurate acquisition of the target user's medical data helps improve the accuracy of subsequent identification of the target user's medical risks based on the medical data. After determining the target user's intended insurance product, an initial underwriting decision tree model corresponding to the type of the intended insurance product is matched, and rules are configured on the initial underwriting decision tree model to obtain a target underwriting decision tree model. Finally, the target underwriting decision tree model is used to perform underwriting identification based on the medical data, and underwriting query results are output according to the identification results of multiple decision trees in the target underwriting decision tree model, thus improving the efficiency and accuracy of risk identification. Attached Figure Description
[0052] Figure 1 This is a flowchart of the medical data risk identification method provided in Embodiment 1 of the present invention.
[0053] Figure 2 This is a structural diagram of the medical data risk identification device provided in Embodiment 2 of the present invention.
[0054] Figure 3 This is a schematic diagram of the structure of the electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0055] To better understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing an embodiment in one alternative implementation and is not intended to be limiting of the invention.
[0057] The medical data risk identification method provided in this embodiment of the invention is executed by an electronic device, and correspondingly, the medical data risk identification device operates in the electronic device.
[0058] Example 1
[0059] Figure 1 This is a flowchart of a medical data risk identification method provided in Embodiment 1 of the present invention. The medical data risk identification method specifically includes the following steps. Depending on different needs, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0060] S11, Receive an underwriting query request for the target user, and determine the target data source based on the underwriting query request.
[0061] Target users refer to individuals who require medical risk identification. Target users can browse insurance products and apply for insurance through the insurance company's terminal, thereby triggering the insurance company's terminal to initiate an underwriting query request to the risk control platform to inquire about the target user's medical risks.
[0062] The risk control platform responds to underwriting query requests from insurance companies' terminals, identifies the target data source for the target user, and accurately obtains the target user's medical data from the target data source, thereby improving the accuracy of identifying the target user's medical risks.
[0063] In an optional implementation, determining the target data source based on the underwriting query request includes:
[0064] Obtain the target user's intended insurance region and authorized institution identifier from the underwriting query request;
[0065] The regional agency mapping table is used to determine the multiple data agency identifiers corresponding to the intended insurance region;
[0066] The target data organization identifier is obtained from the plurality of data organization identifiers based on the authorized organization identifier;
[0067] The data institution corresponding to the target data institution identifier is identified as the target data source.
[0068] The insurance company's terminal displays an insurance purchase page. Target users can enter their intended insurance region in the text input box on the page and select one or more authorized institution identifiers from the authorized institution identifier list. The intended insurance region refers to the region where the target user plans to purchase insurance. The authorized institution identifier indicates which institution the target user authorizes to provide medical data.
[0069] The risk control platform pre-stores a regional agency mapping table, which contains the mapping relationship between insured regions and authorized agency identifiers. One insured region can correspond to multiple authorized agency identifiers, and multiple insured regions can correspond to one authorized agency identifier. Based on the mapping relationship, multiple authorized agency identifiers corresponding to the intended insured region can be determined, thereby matching the authorized agency identifier selected by the target user with the determined multiple authorized agency identifiers. The authorized agency identifier that successfully matches the authorized agency identifier selected by the target user is determined as the target data source.
[0070] The risk control platform initiates a medical data acquisition request to the target data source in order to obtain the target user's medical data.
[0071] In the above optional implementation, determining the target data source through the underwriting query request of the target user helps to accurately obtain the target user's medical data, thereby improving the accuracy of subsequent identification of the target user's medical risks based on the medical data.
[0072] S12, Obtain the target user's medical data from the target data source.
[0073] Requests for medical data can include the target user's identification information, such as name, ID number, social security card number, etc.
[0074] The target data source responds to the medical data acquisition request from the risk control platform and queries the target user's medical data based on the target user's identification information in the medical data acquisition request.
[0075] Medical data may include, but is not limited to: medical records from previous years, examination / test data, medication data, etc.
[0076] In an optional implementation, obtaining the target user's medical data from the target data source includes:
[0077] Obtain the initial medical dataset of the target user from the target data source. Each piece of initial medical data in the initial medical dataset includes patient description information and medical description information.
[0078] The initial medical dataset is sampled to obtain a medical sample set with the same data distribution as the initial medical dataset;
[0079] Medical description values are determined in the medical sample set such that the ratio of the number of initial medical data including the medical description values to the number of medical sample sets is greater than a first preset threshold.
[0080] Obtain patient description values corresponding to the medical description values from the medical sample set, such that the ratio of the number of initial medical data including the patient description values to the number of medical sample sets is greater than a second preset threshold.
[0081] Search the initial medical dataset for initial medical data that includes the medical description value but does not include the patient description value;
[0082] The initial medical data found will be used as the medical data of the target user.
[0083] The initial medical dataset is a massive dataset to be processed. This initial medical dataset includes multiple initial medical data points, each corresponding to a single medical visit and settlement information of a target user.
[0084] The patient description information refers to the relevant description of the target user's own information, which may include the target user's identifier, gender, and age. The target user identifier may be information such as the target user's social security account number, target user's name, or ID card number.
[0085] Among them, medical description information refers to relevant information about the target user's medical visit and the medical items and methods used during the visit, which may include main diagnosis information, drug information, treatment item information and medical service facility information.
[0086] The determined medical description values and the determined patient description values have a strong correlation. Accordingly, the initial medical data that does not satisfy this strong correlation is considered risky medical data. Therefore, in the initial medical dataset, the initial medical data that includes medical description values but does not include patient description values is considered risky medical data.
[0087] The above implementation method can identify high-risk medical data in massive medical datasets. In the identification process, it first determines strongly correlated patient description values and medical description values for the initial medical dataset itself, and then uses the determined strong correlation to judge high-risk medical data in the medical dataset, which has high accuracy.
[0088] S13, obtain the target user's intended insurance product according to the underwriting query request, and match the initial underwriting decision tree model corresponding to the type of the intended insurance product.
[0089] The insurance purchase page displayed on the insurance company's terminal can also receive the intended insurance products entered by the target user. The intended insurance products refer to the products that the target user intends to purchase.
[0090] The types of insurance products that can be considered can be divided into medical insurance, cancer insurance, children's outpatient insurance, chronic disease medical insurance, etc.
[0091] Different types of intended insurance products correspond to different initial underwriting decision models.
[0092] After the target data source returns the target user's medical data to the risk control platform, the risk control platform matches the target user's intended insurance product type to determine the initial underwriting decision tree model. The platform then configures rules based on the determined initial underwriting decision tree model to identify medical risks.
[0093] S14, Configure rules for the initial underwriting decision tree model to obtain the target underwriting decision tree model, which includes multiple decision trees.
[0094] After matching the initial underwriting decision tree model, the risk control platform displays the decision interface corresponding to the initial underwriting decision tree model.
[0095] The risk control platform displays the decision interface corresponding to the initial underwriting decision tree model. In this interface, rules are configured based on the target user's medical data, and the risk control platform then combines these rules to generate the target underwriting decision tree model.
[0096] In an optional implementation, configuring rules on the initial underwriting decision tree model to obtain the target underwriting decision tree model includes:
[0097] Determine whether the target user's medical record identifier can be obtained from the medical data;
[0098] When the medical record identifier of the target user is obtained from the medical data, the number of the medical record identifiers is calculated;
[0099] Configure the number of decision trees according to the stated number;
[0100] Generate the target underwriting decision tree model based on the number of decision trees.
[0101] When no medical record identifier for the target user is obtained from the medical data, it indicates that the target user has no medical records, meaning that the target user has good physical condition and is a low-risk user.
[0102] When the medical record identifier of the target user is obtained from the medical data, it indicates that the target user has a medical record, that is, the target user has poor physical condition and may belong to medium- to high-risk users.
[0103] Each medical record identifier corresponds to a single medical visit. Different correction processes correspond to different medical record identifiers. The more medical record identifiers there are, the more medical visits there are.
[0104] This optional implementation obtains the target underwriting decision tree model by configuring the number of decision trees according to the number of medical record identifiers, ensuring that the number of decision trees in the target underwriting decision tree model corresponds to the number of visits, with one decision tree corresponding to one visit, and one decision tree is used to identify the medical risks of one visit.
[0105] S15, using the target underwriting decision tree model to perform underwriting identification based on the medical data, obtaining the identification result of each decision tree in the target underwriting decision tree model, and outputting the underwriting query result based on the identification results of multiple decision trees.
[0106] The risk control platform transmits the underwriting query results to the insurance company's terminal.
[0107] In an optional implementation, the step of using the target underwriting decision tree model to perform underwriting identification based on the medical data, and obtaining the identification result of each decision tree in the target underwriting decision tree model, includes:
[0108] Obtain the diagnosis and treatment data corresponding to each medical record identifier in the medical data;
[0109] Multiple sets of diagnostic and treatment data are input into the target underwriting decision tree model for underwriting identification;
[0110] Obtain the identification result of each decision tree in the target underwriting decision tree model. The identification result is the disease risk level. The risk control platform determines the underwriting query result by taking the intersection of the results from each decision tree based on the disease risk level.
[0111] In an optional implementation, after outputting the underwriting query result based on the recognition results of multiple decision trees, the method further includes:
[0112] Retrieve incremental data within a predetermined time period;
[0113] Generate an incremental decision tree based on the incremental data;
[0114] Label prediction is performed on the incremental data based on the decision tree in the incremental decision tree and the decision tree in the target underwriting decision tree model to obtain the label prediction result;
[0115] The overall performance of the decision tree in the target underwriting decision tree model and the decision tree in the incremental decision tree is determined based on the label prediction results.
[0116] Based on the overall performance, a predetermined number of decision trees are selected from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree model as the updated target underwriting decision tree model.
[0117] Incremental data refers to the medical data of newly added users within a predetermined time period after the target user.
[0118] The overall performance of each decision tree is determined at least based on the establishment time of each decision tree and the prediction accuracy for the incremental data. The prediction accuracy is obtained by comparing the predicted label results with the actual label results. Specifically, the prediction accuracy is calculated by obtaining the target label prediction results that are identical to the actual label results and then calculating the ratio of the number of target label prediction results to the number of actual label results.
[0119] Multiple sample sets can be extracted with replacement from the incremental data, and multiple incremental decision trees can be generated based on the multiple sample sets. The number of incremental decision trees can range from 10% to 30% of the number of decision trees in the target underwriting decision tree model.
[0120] The predetermined number can be equal to the number of decision trees in the target underwriting decision tree model.
[0121] In an optional implementation, selecting a predetermined number of decision trees from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree includes:
[0122] The prediction accuracy of the decision tree for the incremental data is determined based on the results of the label prediction.
[0123] The time it takes to build the decision tree is used as the weight for the overall performance.
[0124] The prediction accuracy of the incremental data is sorted according to the weights.
[0125] The predetermined number of decision trees are obtained from the decision trees in the target underwriting decision tree model and the sorted prediction accuracy.
[0126] The weight of a decision tree that has been built for a longer period of time is less than the weight of a decision tree that has been built for a shorter period of time.
[0127] This invention determines the target data source by identifying the target user's underwriting query request, thereby accurately obtaining the target user's medical data from the target data source. Accurate acquisition of the target user's medical data helps improve the accuracy of subsequent identification of the target user's medical risks based on the medical data. After determining the target user's intended insurance product, an initial underwriting decision tree model corresponding to the type of the intended insurance product is matched. Rules are then configured on the initial underwriting decision tree model to obtain a target underwriting decision tree model. Finally, the target underwriting decision tree model is used to perform underwriting identification based on the medical data, and underwriting query results are output according to the identification results of multiple decision trees in the target underwriting decision tree model, thus improving the efficiency and accuracy of risk identification.
[0128] Example 2
[0129] Figure 2 This is a structural diagram of the medical data risk identification device provided in Embodiment 2 of the present invention.
[0130] In some embodiments, the medical data risk identification device 20 may include multiple functional modules composed of computer program segments. The computer programs for each program segment in the medical data risk identification device 20 may be stored in the memory of an electronic device and executed by at least one processor to perform (see details). Figure 1 (Description) Function for identifying risks in medical data.
[0131] In this embodiment, the medical data risk identification device 20 can be divided into multiple functional modules according to its functions. These functional modules may include: a receiving module 201, an acquisition module 202, a matching module 203, a configuration module 204, an identification module 205, and an update module 206. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functions of each module will be detailed in subsequent embodiments.
[0132] The receiving module 201 is used to receive an underwriting query request for a target user and determine the target data source based on the underwriting query request.
[0133] Target users refer to individuals who require medical risk identification. Target users can browse insurance products and apply for insurance through the insurance company's terminal, thereby triggering the insurance company's terminal to initiate an underwriting query request to the risk control platform to inquire about the target user's medical risks.
[0134] The risk control platform responds to underwriting query requests from insurance companies' terminals, identifies the target data source for the target user, and accurately obtains the target user's medical data from the target data source, thereby improving the accuracy of identifying the target user's medical risks.
[0135] In an optional implementation, the receiving module 201 determines the target data source based on the underwriting query request by including:
[0136] Obtain the target user's intended insurance region and authorized institution identifier from the underwriting query request;
[0137] The regional agency mapping table is used to determine the multiple data agency identifiers corresponding to the intended insurance region;
[0138] The target data organization identifier is obtained from the plurality of data organization identifiers based on the authorized organization identifier;
[0139] The data institution corresponding to the target data institution identifier is identified as the target data source.
[0140] The insurance company's terminal displays an insurance purchase page. Target users can enter their intended insurance region in the text input box on the page and select one or more authorized institution identifiers from the authorized institution identifier list. The intended insurance region refers to the region where the target user plans to purchase insurance. The authorized institution identifier indicates which institution the target user authorizes to provide medical data.
[0141] The risk control platform pre-stores a regional agency mapping table, which contains the mapping relationship between insured regions and authorized agency identifiers. One insured region can correspond to multiple authorized agency identifiers, and multiple insured regions can correspond to one authorized agency identifier. Based on the mapping relationship, multiple authorized agency identifiers corresponding to the intended insured region can be determined, thereby matching the authorized agency identifier selected by the target user with the determined multiple authorized agency identifiers. The authorized agency identifier that successfully matches the authorized agency identifier selected by the target user is determined as the target data source.
[0142] The risk control platform initiates a medical data acquisition request to the target data source in order to obtain the target user's medical data.
[0143] In the above optional implementation, determining the target data source through the underwriting query request of the target user helps to accurately obtain the target user's medical data, thereby improving the accuracy of subsequent identification of the target user's medical risks based on the medical data.
[0144] The acquisition module 202 is used to acquire the target user's medical data from the target data source.
[0145] Requests for medical data can include the target user's identification information, such as name, ID number, social security card number, etc.
[0146] The target data source responds to the medical data acquisition request from the risk control platform and queries the target user's medical data based on the target user's identification information in the medical data acquisition request.
[0147] Medical data may include, but is not limited to: medical records from previous years, examination / test data, medication data, etc.
[0148] In an optional implementation, the acquisition module 202 acquires the target user's medical data from the target data source, including:
[0149] Obtain the initial medical dataset of the target user from the target data source. Each piece of initial medical data in the initial medical dataset includes patient description information and medical description information.
[0150] The initial medical dataset is sampled to obtain a medical sample set with the same data distribution as the initial medical dataset;
[0151] Medical description values are determined in the medical sample set such that the ratio of the number of initial medical data including the medical description values to the number of medical sample sets is greater than a first preset threshold.
[0152] Obtain patient description values corresponding to the medical description values from the medical sample set, such that the ratio of the number of initial medical data including the patient description values to the number of medical sample sets is greater than a second preset threshold.
[0153] Search the initial medical dataset for initial medical data that includes the medical description value but does not include the patient description value;
[0154] The initial medical data found will be used as the medical data of the target user.
[0155] The initial medical dataset is a massive dataset to be processed. This initial medical dataset includes multiple initial medical data points, each corresponding to a single medical visit and settlement information of a target user.
[0156] The patient description information refers to the relevant description of the target user's own information, which may include the target user's identifier, gender, and age. The target user identifier may be information such as the target user's social security account number, target user's name, or ID card number.
[0157] Among them, medical description information refers to relevant information about the target user's medical visit and the medical items and methods used during the visit, which may include main diagnosis information, drug information, treatment item information and medical service facility information.
[0158] The determined medical description values and the determined patient description values have a strong correlation. Accordingly, the initial medical data that does not satisfy this strong correlation is considered risky medical data. Therefore, in the initial medical dataset, the initial medical data that includes medical description values but does not include patient description values is considered risky medical data.
[0159] The above implementation method can identify high-risk medical data in massive medical datasets. In the identification process, it first determines strongly correlated patient description values and medical description values for the initial medical dataset itself, and then uses the determined strong correlation to judge high-risk medical data in the medical dataset, which has high accuracy.
[0160] The matching module 203 is used to obtain the target user's intended insurance product according to the underwriting query request, and match an initial underwriting decision tree model corresponding to the type of the intended insurance product.
[0161] The insurance purchase page displayed on the insurance company's terminal can also receive the intended insurance products entered by the target user. The intended insurance products refer to the products that the target user intends to purchase.
[0162] The types of insurance products that can be considered can be divided into medical insurance, cancer insurance, children's outpatient insurance, chronic disease medical insurance, etc.
[0163] Different types of intended insurance products correspond to different initial underwriting decision models.
[0164] After the target data source returns the target user's medical data to the risk control platform, the risk control platform matches the target user's intended insurance product type to determine the initial underwriting decision tree model. The platform then configures rules based on the determined initial underwriting decision tree model to identify medical risks.
[0165] The configuration module 204 is used to configure rules for the initial underwriting decision tree model to obtain a target underwriting decision tree model, wherein the target underwriting decision tree model includes multiple decision trees.
[0166] After matching the initial underwriting decision tree model, the risk control platform displays the decision interface corresponding to the initial underwriting decision tree model.
[0167] The risk control platform displays the decision interface corresponding to the initial underwriting decision tree model. In this interface, rules are configured based on the target user's medical data, and the risk control platform then combines these rules to generate the target underwriting decision tree model.
[0168] In an optional implementation, the configuration module 204 configures rules for the initial underwriting decision tree model to obtain the target underwriting decision tree model, including:
[0169] Determine whether the target user's medical record identifier can be obtained from the medical data;
[0170] When the medical record identifier of the target user is obtained from the medical data, the number of the medical record identifiers is calculated;
[0171] Configure the number of decision trees according to the stated number;
[0172] Generate the target underwriting decision tree model based on the number of decision trees.
[0173] When no medical record identifier for the target user is obtained from the medical data, it indicates that the target user has no medical records, meaning that the target user has good physical condition and is a low-risk user.
[0174] When the medical record identifier of the target user is obtained from the medical data, it indicates that the target user has a medical record, that is, the target user has poor physical condition and may belong to medium- to high-risk users.
[0175] Each medical record identifier corresponds to a single medical visit. Different correction processes correspond to different medical record identifiers. The more medical record identifiers there are, the more medical visits there are.
[0176] This optional implementation obtains the target underwriting decision tree model by configuring the number of decision trees according to the number of medical record identifiers, ensuring that the number of decision trees in the target underwriting decision tree model corresponds to the number of visits, with one decision tree corresponding to one visit, and one decision tree is used to identify the medical risks of one visit.
[0177] The identification module 205 is used to perform underwriting identification based on the medical data using the target underwriting decision tree model, obtain the identification result of each decision tree in the target underwriting decision tree model, and output the underwriting query result based on the identification results of multiple decision trees.
[0178] The risk control platform transmits the underwriting query results to the insurance company's terminal.
[0179] In an optional implementation, the identification module 205 uses the target underwriting decision tree model to perform underwriting identification based on the medical data, obtaining the identification result of each decision tree in the target underwriting decision tree model, including:
[0180] Obtain the diagnosis and treatment data corresponding to each medical record identifier in the medical data;
[0181] Multiple sets of diagnostic and treatment data are input into the target underwriting decision tree model for underwriting identification;
[0182] Obtain the identification result of each decision tree in the target underwriting decision tree model.
[0183] The identification result is the disease risk level. The risk control platform uses the disease risk level to find the intersection of the data and determines the underwriting query result.
[0184] In an optional implementation, after outputting the underwriting query result based on the recognition results of multiple decision trees, the update module 206 is used to:
[0185] Retrieve incremental data within a predetermined time period;
[0186] Generate an incremental decision tree based on the incremental data;
[0187] Label prediction is performed on the incremental data based on the decision tree in the incremental decision tree and the decision tree in the target underwriting decision tree model to obtain the label prediction result;
[0188] The overall performance of the decision tree in the target underwriting decision tree model and the decision tree in the incremental decision tree is determined based on the label prediction results.
[0189] Based on the overall performance, a predetermined number of decision trees are selected from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree model as the updated target underwriting decision tree model.
[0190] Incremental data refers to the medical data of newly added users within a predetermined time period after the target user.
[0191] The overall performance of each decision tree is determined at least based on the establishment time of each decision tree and the prediction accuracy for the incremental data. The prediction accuracy is obtained by comparing the predicted label results with the actual label results. Specifically, the prediction accuracy is calculated by obtaining the target label prediction results that are identical to the actual label results and then calculating the ratio of the number of target label prediction results to the number of actual label results.
[0192] Multiple sample sets can be extracted with replacement from the incremental data, and multiple incremental decision trees can be generated based on the multiple sample sets. The number of incremental decision trees can range from 10% to 30% of the number of decision trees in the target underwriting decision tree model.
[0193] The predetermined number can be equal to the number of decision trees in the target underwriting decision tree model.
[0194] In an optional implementation, selecting a predetermined number of decision trees from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree includes:
[0195] The prediction accuracy of the decision tree for the incremental data is determined based on the results of the label prediction.
[0196] The time it takes to build the decision tree is used as the weight for the overall performance.
[0197] The prediction accuracy of the incremental data is sorted according to the weights.
[0198] The predetermined number of decision trees are obtained from the decision trees in the target underwriting decision tree model and the sorted prediction accuracy.
[0199] The weight of a decision tree that has been built for a longer period of time is less than the weight of a decision tree that has been built for a shorter period of time.
[0200] This invention determines the target data source by identifying the target user's underwriting query request, thereby accurately obtaining the target user's medical data from the target data source. Accurate acquisition of the target user's medical data helps improve the accuracy of subsequent identification of the target user's medical risks based on the medical data. After determining the target user's intended insurance product, an initial underwriting decision tree model corresponding to the type of the intended insurance product is matched. Rules are then configured on the initial underwriting decision tree model to obtain a target underwriting decision tree model. Finally, the target underwriting decision tree model is used to perform underwriting identification based on the medical data, and underwriting query results are output according to the identification results of multiple decision trees in the target underwriting decision tree model, thus improving the efficiency and accuracy of risk identification.
[0201] Example 3
[0202] This embodiment provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps described in the above-described medical data risk identification method embodiment, for example... Figure 1 S11-S15 as shown:
[0203] S11, Receive an underwriting query request for the target user, and determine the target data source based on the underwriting query request;
[0204] S12, Obtain the target user's medical data from the target data source;
[0205] S13, obtain the target user's intended insurance product according to the underwriting query request, and match the initial underwriting decision tree model corresponding to the type of the intended insurance product;
[0206] S14, Configure rules for the initial underwriting decision tree model to obtain the target underwriting decision tree model, wherein the target underwriting decision tree model includes multiple decision trees;
[0207] S15, using the target underwriting decision tree model to perform underwriting identification based on the medical data, obtaining the identification result of each decision tree in the target underwriting decision tree model, and outputting the underwriting query result based on the identification results of multiple decision trees.
[0208] Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 Modules 201-205 in the middle:
[0209] The receiving module 201 is used to receive an underwriting query request for a target user and determine the target data source based on the underwriting query request.
[0210] The acquisition module 202 is used to acquire the medical data of the target user from the target data source;
[0211] The matching module 203 is used to obtain the target user's intended insurance product according to the underwriting query request, and match an initial underwriting decision tree model corresponding to the type of the intended insurance product;
[0212] The configuration module 204 is used to configure rules for the initial underwriting decision tree model to obtain a target underwriting decision tree model, wherein the target underwriting decision tree model includes multiple decision trees;
[0213] The identification module 205 is used to perform underwriting identification based on the medical data using the target underwriting decision tree model, obtain the identification result of each decision tree in the target underwriting decision tree model, and output the underwriting query result based on the identification results of multiple decision trees.
[0214] When the computer program is executed by the processor, it implements the update module 206 in the above device embodiment. For details, please refer to Embodiment 2 and its related description.
[0215] Example 4
[0216] See Figure 3 The diagram shown is a structural schematic of an electronic device provided in Embodiment 3 of the present invention. In a preferred embodiment of the present invention, the electronic device 3 includes a memory 31, at least one processor 32, at least one communication bus 33, and a transceiver 34.
[0217] Those skilled in the art should understand that Figure 3The structure of the electronic device shown does not constitute a limitation of the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 3 may also include more or fewer other hardware or software than shown, or different component arrangements.
[0218] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital processors, and embedded devices. The electronic device 3 may also include client devices, including, but not limited to, any electronic product capable of human-computer interaction with a client via a keyboard, mouse, remote control, touchpad, or voice control device, such as personal computers, tablets, smartphones, and digital cameras.
[0219] It should be noted that the electronic device 3 is merely an example. Other existing or future electronic products that are suitable for this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0220] In some embodiments, the memory 31 stores a computer program that, when executed by the at least one processor 32, implements all or part of the steps in the medical data risk identification method described above. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0221] Furthermore, the computer-readable storage medium may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system, at least one application required for a function, etc.; and the data storage area may store data created based on the use of blockchain nodes, etc.
[0222] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0223] In some embodiments, the at least one processor 32 is the control unit of the electronic device 3, connecting various components of the electronic device 3 via various interfaces and lines. It executes programs or modules stored in the memory 31 and calls data stored in the memory 31 to perform various functions and process data. For example, when the at least one processor 32 executes a computer program stored in the memory, it implements all or part of the steps of the medical data risk identification method described in this embodiment of the invention; or it implements all or part of the functions of the medical data risk identification device. The at least one processor 32 may be composed of integrated circuits, such as a single-packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.
[0224] In some embodiments, the at least one communication bus 33 is configured to enable communication between the memory 31 and the at least one processor 32, etc.
[0225] Although not shown, the electronic device 3 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 32 via a power management device, thereby enabling functions such as charging, discharging, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 3 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0226] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) or processor to execute portions of the methods described in the various embodiments of the present invention.
[0227] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0228] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0229] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0230] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements, and the singular does not exclude the plural. Multiple elements or devices recited in the specification may also be implemented by a single element or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
[0231] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for identifying risks in medical data, characterized in that, The method includes: Receive an underwriting query request for a target user, and obtain the target user's intended insurance region and authorized agency identifier from the underwriting query request; determine multiple data agency identifiers corresponding to the intended insurance region based on a region-agency mapping table; obtain the target data agency identifier from the multiple data agency identifiers based on the authorized agency identifier; and determine the data agency corresponding to the target data agency identifier as the target data source. Obtain the target user's medical data from the target data source; The target user's intended insurance products are obtained based on the underwriting query request, and an initial underwriting decision tree model corresponding to the type of the intended insurance product is matched. Different types of intended insurance products correspond to different initial underwriting decision models. The medical record identifier of the target user is obtained from the medical data. Based on the number of medical record identifiers, the rules of the initial underwriting decision tree model are configured to obtain the target underwriting decision tree model. The target underwriting decision tree model includes multiple decision trees, and the number of decision trees corresponds to the number of medical record identifiers. The target underwriting decision tree model is used to perform underwriting identification based on the medical data, and the identification results of each decision tree in the target underwriting decision tree model are obtained. The underwriting query results are then output based on the identification results of multiple decision trees.
2. The medical data risk identification method as described in claim 1, characterized in that, The step of obtaining the target user's medical data from the target data source includes: Obtain the initial medical dataset of the target user from the target data source. Each piece of initial medical data in the initial medical dataset includes patient description information and medical description information. The initial medical dataset is sampled to obtain a medical sample set with the same data distribution as the initial medical dataset; Medical description values are determined in the medical sample set such that the ratio of the number of initial medical data including the medical description values to the number of medical sample sets is greater than a first preset threshold. Obtain patient description values corresponding to the medical description values from the medical sample set, such that the ratio of the number of initial medical data including the patient description values to the number of medical sample sets is greater than a second preset threshold. Search the initial medical dataset for initial medical data that includes the medical description value but does not include the patient description value; The initial medical data found will be used as the medical data of the target user.
3. The medical data risk identification method as described in claim 1, characterized in that, The process involves obtaining the target user's medical record identifier from the medical data, configuring rules for the initial underwriting decision tree model based on the number of medical record identifiers, and obtaining a target underwriting decision tree model. The target underwriting decision tree model includes multiple decision trees, including: Determine whether the target user's medical record identifier can be obtained from the medical data; When the medical record identifier of the target user is obtained from the medical data, the number of the medical record identifiers is calculated; Configure the number of decision trees according to the stated number; Generate the target underwriting decision tree model based on the number of decision trees.
4. The medical data risk identification method as described in any one of claims 1 to 3, characterized in that, The process of using the target underwriting decision tree model to perform underwriting identification based on the medical data, and obtaining the identification result of each decision tree in the target underwriting decision tree model, includes: Obtain the diagnosis and treatment data corresponding to each medical record identifier in the medical data; Multiple sets of diagnostic and treatment data are input into the target underwriting decision tree model for underwriting identification; Obtain the identification result of each decision tree in the target underwriting decision tree model.
5. The medical data risk identification method as described in claim 4, characterized in that, After outputting the underwriting query result based on the recognition results of multiple decision trees, the method further includes: Retrieve incremental data within a predetermined time period; Generate an incremental decision tree based on the incremental data; Label prediction is performed on the incremental data based on the decision tree in the incremental decision tree and the decision tree in the target underwriting decision tree model to obtain the label prediction result; The overall performance of the decision tree in the target underwriting decision tree model and the decision tree in the incremental decision tree is determined based on the label prediction results. Based on the overall performance, a predetermined number of decision trees are selected from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree model as the updated target underwriting decision tree model.
6. The medical data risk identification method as described in claim 5, characterized in that, The step of selecting a predetermined number of decision trees from the decision trees in the target underwriting decision tree model and the decision trees in the incremental decision tree includes: The prediction accuracy of the decision tree for the incremental data is determined based on the results of the label prediction. The time it takes to build the decision tree is used as the weight for the overall performance. The prediction accuracy of the incremental data is sorted according to the weights. The predetermined number of decision trees are obtained from the decision trees in the target underwriting decision tree model and the sorted prediction accuracy.
7. A medical data risk identification device, characterized in that, The device includes: The receiving module is configured to receive an underwriting query request for a target user, obtain the target user's intended insurance region and authorized agency identifier from the underwriting query request; determine multiple data agency identifiers corresponding to the intended insurance region based on a region-agency mapping table; obtain a target data agency identifier from the multiple data agency identifiers based on the authorized agency identifier; and determine the data agency corresponding to the target data agency identifier as the target data source. The acquisition module is used to acquire the target user's medical data from the target data source; The matching module is used to obtain the target user's intended insurance products according to the underwriting query request, and match them with an initial underwriting decision tree model corresponding to the type of the intended insurance product. Different types of intended insurance products correspond to different initial underwriting decision models. The configuration module is used to obtain the medical record identifier of the target user from the medical data, configure the rules of the initial underwriting decision tree model based on the number of medical record identifiers, and obtain the target underwriting decision tree model. The target underwriting decision tree model includes multiple decision trees, and the number of decision trees corresponds to the number of medical record identifiers. The identification module is used to perform underwriting identification based on the medical data using the target underwriting decision tree model, obtain the identification result of each decision tree in the target underwriting decision tree model, and output the underwriting query result based on the identification results of multiple decision trees.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the processor is configured to implement the medical data risk identification method as described in any one of claims 1 to 6 when executing a computer program stored in the memory.
9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the medical data risk identification method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Medical data auditing method and device, electronic equipment and computer readable medium
CN111179096A