A health insurance automatic quotation system based on big data
By collecting multi-source heterogeneous data, dynamically assigning weights, and aligning them in time and space, combined with risk transmission prediction, a tiered premium mechanism is generated. This solves the problems of distorted health profiles and rigid risk assessments in health insurance, and achieves dynamic matching between premiums and user risk trajectories, thereby improving the personalization and accuracy of health insurance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PICC HEALTH INSURANCE CO LTD
- Filing Date
- 2025-06-25
- Publication Date
- 2026-05-08
AI Technical Summary
The current health insurance industry lacks the ability to integrate and process multi-source heterogeneous health big data, resulting in distorted health profiles, rigid risk assessments, and an inability to respond to individual health trajectories in real time, leading to adverse selection and reduced opportunities for missed claims.
User behavior data is acquired through a multi-source heterogeneous data acquisition module, the dynamic weight allocation module adjusts the weight of data dimensions, the spatiotemporal alignment module unifies the data time frame, the risk transmission prediction module predicts the disease development path, and a tiered premium mechanism is generated to dynamically adjust the premium to match the user's risk trajectory.
It achieves personalized and accurate health risk assessment, reduces adverse selection risk, optimizes the prediction of claims probability, dynamically matches premiums with users' risk trajectories, and breaks through the bottleneck of traditional static pricing.
Smart Images

Figure CN120807087B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, specifically to an automated health insurance quotation system based on big data. Background Technology
[0002] Currently, the health insurance industry generally adopts a fixed premium pricing model based on static actuarial tables. Its risk assessment heavily relies on limited historical data from the initial stage of insurance purchase (such as physical examination reports and past medical history), and cannot incorporate dynamic indicators such as sleep, exercise, and environmental exposure in real time. The existing technological bottlenecks are mainly reflected in the following aspects: First, there is a lack of ability to integrate and process multi-source heterogeneous health big data. The second-level physiological indicators, hour-level activity data and day-level environmental information from wearable devices, smart homes, and mobile applications are separated in terms of time and space. Traditional systems have difficulty aligning the time granularity and quantifying the risk contribution weight of different data sources, resulting in distorted profiles.
[0003] Secondly, the risk transmission mechanism is rigid—once the premium is determined, it remains unchanged for a long period. This makes it impossible to capture potential signals of accelerated disease deterioration (such as regular heart rate abnormalities) behind sustained low-entropy (high-stability) behavioral patterns, nor can it translate predicted risk evolution paths into preventative actuarial interventions. This disconnect between static pricing and dynamic health trajectories not only leads to adverse selection among high-risk groups but also causes insurance companies to miss the window of opportunity to reduce claims risk through behavioral interventions, thus hindering product innovation and market sustainability. Summary of the Invention
[0004] The purpose of this invention is to provide an automatic health insurance pricing system based on big data to solve the problems mentioned in the background. The specific technologies include how to achieve dynamic weight allocation and spatiotemporal unification of multi-dimensional user behavioral data to solve the problem of health profile distortion caused by differences in the value assessment of heterogeneous data and inconsistent time granularity; and how to integrate behavioral stability analysis and risk transmission prediction to solve the problem that traditional premium models cannot respond to the dynamic matching of individual real-time risk trajectories.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] This big data-based automated health insurance pricing system includes a multi-source heterogeneous data acquisition module, a dynamic health profile construction module, and a risk transmission prediction module, wherein:
[0007] The multi-source heterogeneous data acquisition module is used to acquire user behavior data in real time, including sleep depth data, health behavior data, and user living environment data, providing a complete data foundation for subsequent analysis.
[0008] The dynamic weight allocation unit in the dynamic health profile construction module calculates and dynamically adjusts the weight coefficients of different behavioral data dimensions in real time based on individual user characteristics, data timeliness, and correlation with specific disease risks. This solves the problem of lack of personalization caused by static weights. For example, the heart rate data of users with a history of cardiovascular disease has a higher weight than dietary habit data.
[0009] The spatiotemporal alignment and vector weighting unit in the dynamic health profile construction module uses aggregation technology to unify user behavior data at different frequencies and time points, such as second-level and hour-level, into a daily-dimensional time frame, forming a daily-dimensional user behavior vector; and applies dynamic weight coefficients to weight each data point in the vector to generate a weighted daily-dimensional user behavior vector, eliminating spatiotemporal differences between multi-source data.
[0010] The risk profile and entropy value generation unit in the dynamic health profile construction module inputs the weighted daily dimension vector into the disease risk mapping model to calculate the risk probability value of the preset disease dimension; the risk probability values of each disease dimension are weighted and summed and processed by the scaling function to generate a real-time dynamic health risk score.
[0011] In addition, the risk profiling and entropy generation unit collects weighted daily user behavior vectors over a continuous preset time period to form a time series data window; calculates the statistical variation index of each behavioral feature within the window; calculates the average of all feature variation indices to obtain a preliminary fluctuation measure; and inputs this preliminary fluctuation measure into the information entropy formula to generate a health behavior entropy value. This entropy value directly quantifies behavioral stability; low entropy indicates stable behavior, while high entropy indicates disordered behavior.
[0012] The risk transmission prediction module monitors the entropy value sequence of health behaviors in real time through a time series prediction model. It automatically detects segments where the entropy value is lower than the preset stability threshold at multiple consecutive time points, marks them as abnormal duration periods, and accurately records their duration and timestamp.
[0013] The risk transmission prediction module accesses a disease progression rate knowledge base, which stores disease deterioration process models and the mapping relationship between abnormal duration and disease progression acceleration coefficient. Combining the abnormal duration and real-time dynamic health risk score, it predicts the estimated time point for reaching each risk escalation node and outputs a personalized risk evolution path map.
[0014] The risk transmission prediction module extracts the time interval data of adjacent risk escalation nodes from the personalized risk evolution path map, and converts it into an actuarial factor that reflects the degree of risk acceleration through an actuarial conversion algorithm. The shorter the time interval, the larger the actuarial factor value.
[0015] The risk transmission prediction module applies actuarial factors to the pricing engine to perform three steps: dividing the premium tier nodes according to the predicted risk level and arrival time; adjusting the base premium rate by applying actuarial factors within the premium period corresponding to each node; generating a tiered future premium schedule that includes time nodes and corresponding new premium amounts; and when new data is input, the risk transmission prediction module recalculates the actuarial factors and risk evolution path diagram, dynamically updating the tiered premium gradient to ensure that it always matches the user's real-time risk trajectory.
[0016] Compared with the prior art, the beneficial effects of the present invention are:
[0017] By dynamically weighting and aligning multi-source data in time and space, the personalization and accuracy of health risk assessment are improved; by quantifying the stability of behavioral patterns using health behavior entropy values and combining abnormal duration identification with disease progression rate knowledge base modeling, the predictability of behavioral risk transmission paths is achieved; a tiered premium mechanism based on actuarial factor generation enables dynamic and accurate matching of premium levels with users' real-time risk trajectories, breaking through the bottleneck of static pricing in traditional health insurance; and a closed loop of "data collection → risk assessment → transmission prediction → premium output" is formed, effectively reducing adverse selection risk and optimizing claims probability prediction. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall modules of the present invention;
[0019] Figure 2 This is a schematic diagram of the dynamic health profile construction module unit of the present invention.
[0020] In the diagram: 100, Multi-source heterogeneous data acquisition module; 200, Dynamic health profile construction module; 201, Dynamic weight allocation unit; 202, Spatiotemporal alignment and vector weighting unit; 203, Risk profile and entropy generation unit; 300, Risk transmission prediction module. Detailed Implementation
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Next, please refer to Figure 1 The present invention provides a technical solution: an automatic health insurance quotation system based on big data, including a multi-source heterogeneous data acquisition module 100, a dynamic health profile construction module 200, and a risk transmission prediction module 300.
[0023] The multi-source heterogeneous data acquisition module 100 is the core data entry point of this health insurance automatic quotation system. Its core function is to collect users' daily behavior data in real time and continuously from multiple channels with different sources and structures. The user behavior data acquired by this module specifically covers the following three key dimensions:
[0024] Sleep depth data: Records specific indicators of a user's sleep quality, such as sleep duration, the ratio of deep sleep to light sleep, and the number of sleep interruptions;
[0025] Health behavior data: Information reflecting users' proactive health management activities, including but not limited to exercise volume (steps, exercise duration, intensity), dietary habits (calorie intake, nutritional structure), heart rate, blood pressure and other physiological indicator monitoring data;
[0026] User living environment data: Information describing the external environment in which the user is located, such as air quality index (PM2.5, PM10), ambient temperature and humidity, noise level, climate characteristics of the residential area, etc.
[0027] Please see Figure 2 The dynamic weight allocation unit 201 in the dynamic health profile construction module 200 first uses the dynamic weight allocation method to process the collected user behavior data. The core of the dynamic weight allocation method is to calculate and dynamically adjust the weight coefficients of different behavioral data dimensions (such as sleep depth, exercise volume, and environmental PM2.5) in real time based on each user's individual characteristics (such as age, history of underlying diseases, and current health status) as well as the timeliness of the data itself and its correlation with specific disease risks.
[0028] This process means that for different users, or for the same user at different times, the degree of influence (weight) of the same type of data (such as sleep quality) on the overall health risk score is not fixed, but can be flexibly adapted to changes in individual status and differences in data value, ensuring that the weight allocation is more personalized and timely. For example, for users with a history of cardiovascular disease, recent abnormal heart rate data may be given a higher weight than their dietary habit data.
[0029] To address the issues of inconsistent time granularity (such as heart rate at the second level, steps at the hour level, sleep reports at the day level, and even physical examination data at the weekly / monthly level) and spatial source differences in user behavior data acquired from multi-source heterogeneous acquisition modules, the spatiotemporal alignment and vector weighting unit 202 in the dynamic health profile construction module 200 performs crucial spatiotemporal alignment processing. This spatiotemporal alignment processing unifies and standardizes all these data collected at different frequencies and time points into a standard "daily" dimension time frame through aggregation technology, forming a structured "daily dimension user behavior vector." Each vector represents the summary or representative value of the user's various behavioral data within a day.
[0030] Subsequently, the spatiotemporal alignment and vector weighting unit 202 applies the weight coefficients of each data dimension obtained by dynamic calculation to the aligned daily dimension vector, performs weighted calculation on each data point in the vector, and finally generates a weighted daily dimension user behavior vector that can more accurately reflect the actual contribution of each behavioral data to the health risk of the day, providing a standardized and value-quantifiable input for subsequent risk probability calculation.
[0031] The risk profile and entropy generation unit 203 in the dynamic health profile construction module 200 inputs the weighted daily user behavior vector into the disease risk mapping model. This disease risk mapping model is a pre-trained multi-task learning model or an ensemble model containing multiple independent sub-models. Internally, for each preset specific disease dimension (such as cardiovascular disease, diabetes, respiratory diseases, etc.), the disease risk mapping model uses a subset of features related to the disease (for example, for cardiovascular disease, it may focus on the proportion of deep sleep in sleep depth, exercise intensity, resting heart rate, etc.) and performs real-time inference through its specific calculation logic (such as logistic regression, gradient boosting tree, or forward propagation calculation of neural networks). The output result is the predicted probability value of the user developing the disease in a specific time period in the future (e.g., the next year) under the current daily dimension behavior data. The risk probability values of all preset disease dimensions are calculated in parallel or sequentially in this step.
[0032] After obtaining the risk probability values for all preset disease dimensions, the risk profiling and entropy generation unit 203 performs an aggregation operation to generate a single comprehensive health risk score. The aggregation process adopts a weighted summation method, assigning an importance weight reflecting the relative severity of the disease or its impact on overall health to the risk probability value of each disease dimension (these weights can be preset static weights or dynamically fine-tuned based on the user's basic information); the weighted risk probability values of all disease dimensions are added together to obtain a preliminary aggregated value; in order to standardize the score to an easily understandable range (such as 0-100 points or 0-10 points), this preliminary aggregated value is processed by a scaling function (such as a linear transformation or a sigmoid function); the final output real-time dynamic health risk score is a numerical value that intuitively reflects the user's current overall health risk level, and this score is dynamically updated as new input vectors are processed daily.
[0033] The risk profiling and entropy generation unit 203 simultaneously focuses on the stability of user health behavior. It collects and stores weighted daily-dimensional health behavior vectors from multiple consecutive preset time periods (e.g., 30 or 90 consecutive days) to form a time-series data window. For the vector sequence within this time window, fluctuation feature extraction is performed. The specific method is...
[0034] First, calculate the statistical variation index (usually the coefficient of variation or standard deviation divided by the mean) of each behavioral feature (such as sleep duration, steps, and environmental PM2.5 exposure) in the daily vector within the time window.
[0035] Then, the average or weighted average of all characteristic variation indicators is calculated to obtain a preliminary fluctuation measure. In order to more accurately characterize the disorder or uncertainty of behavioral patterns, this preliminary fluctuation measure is input into an information entropy calculation formula (such as the Shannon entropy formula). The calculated entropy value is the health behavior entropy value. The higher the entropy value, the greater the fluctuation and the more unstable the user's health behavior pattern has been in the past period. The lower the entropy value, the more regular and stable the behavior pattern is. This entropy value is output as a key indicator for evaluating the consistency of user behavior.
[0036] The risk transmission prediction module 300 continuously receives health behavior entropy values output by the risk profiling and entropy generation unit 203, forming an entropy time series. This entropy time series is input into a pre-trained time series prediction model (e.g., a Long Short-Term Memory network based on a recurrent neural network, or Prophet based on an additive regression model). This time series prediction model not only predicts future entropy trends but, more importantly, monitors sequence changes in real time. The module internally sets a preset stability threshold (e.g., entropy < 2.0, indicating extremely stable behavior). Through model state or post-processing rules (e.g., state machine or sliding window counter), it automatically detects segments where the entropy values at multiple consecutive time points (e.g., 30 consecutive days) are all below the threshold. Once such segments are detected, they are marked as abnormal duration periods, and the duration (e.g., the number of consecutive low-entropy days), start and end timestamps are accurately recorded. The key to this process lies in the model's sequence pattern recognition ability and the scientific setting of the stability threshold.
[0037] After identifying an abnormal duration, the risk transmission prediction module 300 accesses a pre-built and continuously updated disease progression rate knowledge base. This knowledge base is a mapping relation database or rule engine, whose core stores models of the deterioration process of diseases in different dimensions (such as hypertension and type II diabetes). These models quantify the influence coefficient (or acceleration factor) of the duration of abnormal behavior (i.e., the low-entropy stable period) at a specific duration (e.g., 3 months, 6 months, 1 year) on the progression rate of various diseases. Combining the currently identified specific duration of abnormal behavior with the user's current real-time dynamic health risk score (which implies the current overall risk level), the module applies the mapping rules in the knowledge base (which may involve interpolation, lookup, or formula calculation). Through extrapolation and calculation, it predicts the estimated time points (time axis) from the current health state to the point where key disease milestones (such as pathological indicators exceeding the standard, the appearance of symptoms, clinical diagnosis, and other risk escalation nodes) are reached while maintaining this abnormal behavior pattern. Finally, it integrates all risk escalation nodes and their corresponding time points and outputs them as a visualized personalized risk evolution path diagram in the form of a timeline chart. This diagram intuitively shows the expected trajectory of future health risks over time.
[0038] The risk transmission prediction module 300 extracts key information from the generated risk evolution path diagram, namely the time interval between each adjacent risk escalation node (e.g., it is estimated that it will take 12 months to escalate from "current low risk" to "medium risk"; and another 8 months to escalate from "medium risk" to "high risk"). This time interval data is input into an actuarial conversion algorithm, which is based on actuarial principles and statistical models and designed to quantify the impact of the risk evolution speed. The core logic of the algorithm is to identify "shorter escalation time intervals" as a sign of "higher risk acceleration". Through the actuarial factor calculation formula (e.g., actuarial factor = benchmark factor * (standard escalation time / actual predicted escalation time)^k, where k is an empirical constant), the time interval data (representing the risk evolution rate) is converted into one or more numerical actuarial adjustment factors. The actuarial factor (usually greater than 1) is calculated, and the higher the value, the more severe the risk deterioration within the same time period, and the faster the user's future insurance claim risk increases.
[0039] The risk transmission prediction module 300 applies the calculated actuarial factors to the user's insurance product pricing engine. The pricing engine is not a fixed rate, but is designed to be dynamically responsive. Based on the actuarial factors, the pricing engine performs the following operations:
[0040] Determine risk level segments: Based on the predicted risk level (such as low, medium, high, and extremely high) in the risk evolution path diagram and its expected arrival time, divide the risk into multiple premium ladder nodes (time points).
[0041] Calculate tiered premiums: Within the premium period corresponding to each premium tier node (e.g., the current to upgrade point A is the first tier, and the period from upgrade point A to upgrade point B is the second tier), the actuarial factor is applied to adjust the base premium rate for that tier. The actuarial factor is mainly used to amplify the base premium rate of subsequent higher risk tiers (e.g., the base premium rate of the second tier * actuarial factor = the user's actual premium rate for the second tier).
[0042] Generate gradient structure: The engine outputs a step-like future premium schedule, which includes specific time nodes and corresponding new premium amounts; each node represents the estimated time point of risk escalation, after which the premium jumps to a new higher step;
[0043] Dynamic synchronization: The entire process is dynamic. As time goes by and new data is input (new entropy values, new health risk scores), the risk transmission prediction module 300 will continuously repeat the operation of the pricing engine. Once a new abnormal duration or a significant change in risk status is detected, the actuarial factors and prediction path map will be updated, which will cause the tiered premium gradient to be dynamically recalculated and adjusted to ensure that it always matches the user's latest predicted risk trajectory. This achieves a refined and dynamic match between premiums and individual risk status in the time dimension.
[0044] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A health insurance automatic quotation system based on big data, characterized in that, It includes a multi-source heterogeneous data acquisition module (100), a dynamic health profile construction module (200), and a risk transmission prediction module (300), wherein: The multi-source heterogeneous data acquisition module (100) is used to acquire user behavior data in real time, including sleep depth data, health behavior data and user living environment data. The dynamic health profile construction module (200) uses a dynamic weight allocation method to calculate dynamic weight coefficients for different data in user behavior data; and the dynamic health profile construction module (200) performs spatiotemporal alignment processing on user behavior data to unify data of different time granularities into a daily dimension user behavior vector, and generates a weighted daily dimension user behavior vector based on the dynamic weight coefficients. The dynamic health profile construction module (200) calculates the risk probability value of each disease dimension based on the weighted daily dimension health behavior vector through the disease risk mapping model, and aggregates the risk probability value into a real-time dynamic health risk score. At the same time, it extracts the fluctuation characteristics of the vector in a continuous preset time period to generate a health behavior entropy value. The risk transmission prediction module (300) inputs the health behavior entropy value into the time series prediction model to identify the abnormal duration period when the health behavior entropy value is continuously lower than a preset threshold; based on the mapping relationship between the abnormal duration period and the disease development rate, it outputs a personalized risk evolution path diagram; it converts the risk escalation time nodes in the path diagram into actuarial factors, and dynamically generates a step-by-step premium gradient that matches the user's real-time risk trajectory based on the actuarial factors. The dynamic health profile construction module (200) includes a dynamic weight allocation unit (201), which is used to calculate and dynamically adjust the weight coefficients of different behavioral data dimensions in real time based on the user's individual characteristics, data timeliness and correlation with specific disease risks. The dynamic health profile construction module (200) includes a risk profile and entropy value generation unit (203). The risk profile and entropy value generation unit (203) is used to input the weighted daily dimension user behavior vector into the disease risk mapping model, calculate the risk probability value of each preset disease dimension, and perform weighted summation of the risk probability values of each disease dimension and process them through a scaling function to generate a real-time dynamic health risk score. The risk profiling and entropy generation unit (203) forms a time series data window by collecting weighted daily dimension user behavior vectors over a continuous preset time period, calculates the statistical variation index of each behavioral feature within the window, calculates the average value of all feature variation indexes to obtain a preliminary fluctuation measure, and inputs the preliminary fluctuation measure into the information entropy formula to generate a health behavior entropy value. The risk transmission prediction module (300) accesses a disease development rate knowledge base, which stores disease deterioration process models of different dimensions. The model quantifies the impact coefficient of abnormal duration of a specific duration on the disease development rate. Combining the duration of the abnormal duration and the real-time dynamic health risk score, it predicts the estimated time point for reaching each risk escalation node and generates a personalized risk evolution path map. The risk transmission prediction module (300) extracts the time interval data between adjacent risk escalation nodes from the personalized risk evolution path map, and converts the time interval data into an actuarial factor reflecting the degree of risk acceleration through an actuarial conversion algorithm.
2. The health insurance automatic quotation system based on big data according to claim 1, characterized in that, The dynamic health profile construction module (200) includes a spatiotemporal alignment and vector weighting unit (202). The spatiotemporal alignment and vector weighting unit (202) is used to unify user behavior data collected at different frequencies and at different time points into a daily time frame through aggregation technology to form a daily user behavior vector; and to apply dynamic weight coefficients to perform weighted calculations on the data in the vector to generate a weighted daily user behavior vector.
3. The health insurance automatic quotation system based on big data according to claim 1, characterized in that, The risk transmission prediction module (300) monitors the health behavior entropy value sequence in real time through a time series prediction model, automatically detects segments where the health behavior entropy value is lower than a preset stability threshold at multiple consecutive time points, marks them as abnormal duration periods, and records their duration and timestamp.
4. The health insurance automatic quotation system based on big data according to claim 1, characterized in that, The risk transmission prediction module (300) applies actuarial factors to the pricing engine and performs the following operations: The premium tiers are defined based on the predicted risk level and its expected arrival time. Within the premium period corresponding to each premium tier, actuarial factors are applied to adjust the base premium rate and calculate the tiered premium. A future premium schedule with tiered changes including time nodes and corresponding new premium amounts is generated.
5. The health insurance automatic quotation system based on big data according to claim 1, characterized in that, The risk transmission prediction module (300) recalculates actuarial factors and personalized risk evolution path diagrams when new data is input, and dynamically updates the tiered premium gradient to ensure that it matches the user's real-time risk trajectory.
Citation Information
Patent Citations
Human body real-time monitoring health data design life insurance scheme system
CN117408823A
Health management method based on behavior characteristics
CN117524471A