Personalized insurance product recommendation method and device based on LSTM technology
By building a personalized insurance product recommendation system using LSTM technology, the problem of inaccurate recommendation results in existing systems is solved. It realizes personalized recommendations based on customer health data and behavior, improves recommendation accuracy and customer satisfaction, and reduces operating costs.
Patent Information
- Application Number
- CN202511073250.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing insurance product recommendation systems are based on simple rules or statistical models, resulting in low accuracy of recommendation results and a failure to fully integrate customers' health data, lifestyle data, and historical insurance data.
This paper adopts a personalized insurance product recommendation method based on LSTM technology. By collecting customers' raw data, preprocessing and structuring it, a health profile is constructed. Then, by combining machine learning algorithms and dynamic recommendation algorithms, the health risks and potential needs of customers are analyzed to construct personalized insurance product recommendation schemes.
It improved the accuracy of insurance product recommendations, reduced invalid recommendations, enhanced customer experience and satisfaction, lowered the operating and risk costs of insurance companies, increased customers' enthusiasm for health management, and achieved a two-way match between customers and insurance companies.
Smart Images

Figure CN120975930A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent recommendation, and in particular to a method, apparatus, device, and medium for personalized insurance product recommendation based on LSTM technology. Background Technology
[0002] In the development of business systems for fintech and healthcare, the advancements in artificial intelligence, big data, and recommendation systems have led some companies to explore converting rejected policyholders into potential customers through referral systems. Existing systems typically rely on simple rules or statistical models to integrate and analyze customer health data, lifestyle habits, and historical insurance data. However, the lack of comprehensive data integration and analysis results in low accuracy of the recommendations. Summary of the Invention
[0003] This invention provides a personalized insurance product recommendation method, apparatus, computer equipment, and medium based on LSTM technology to solve the technical problem that the accuracy of recommendation results is low in existing systems based on simple rule recommendation strategies.
[0004] Firstly, a personalized insurance product recommendation method based on LSTM technology is provided, including:
[0005] Collect raw customer data and preprocess the raw data, wherein the raw customer data includes at least: the customer's health data, lifestyle data and historical insurance data;
[0006] The preprocessed raw data is structured based on LSTM technology, and then the raw data and LSTM processing results are stored in a distributed storage system.
[0007] Extract customer health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset;
[0008] Based on machine learning algorithms, data features of customers' health data, lifestyle data, and historical insurance data are extracted from the customer dataset, and the data features are trained to construct a health profile of the customer.
[0009] LSTM technology is used to quickly search and obtain real-time health change trends and real-time behavior data of customers, and to regularly update the health profile of customers.
[0010] Collect reasons for rejection and combine them with updated customer health profiles to analyze customer health risks and potential needs;
[0011] Based on the dynamic recommendation algorithm, the updated health portrait of the customer, the reason for refusal, the health risk of the customer and the potential demand are combined to construct a personalized insurance product recommendation scheme.
[0012] In a second aspect, a personalized insurance product recommendation device based on LSTM technology is provided, comprising:
[0013] The acquisition module is configured to acquire original data of a customer and pre-process the original data, wherein the original data of the customer at least includes health data, living habit data and historical insurance data of the customer.
[0014] The storage module is configured to perform structured processing on the pre-processed original data based on the LSTM technology, and store the original data and the LSTM processing result into a distributed storage system.
[0015] The extraction module is configured to extract the health data, the living habit data and the historical insurance data of the customer from the distributed storage system, and integrate the data into a complete customer data set.
[0016] The training module is configured to extract data features of the health data, the living habit data and the historical insurance data of the customer in the customer data set based on a machine learning algorithm, and train the data features to construct a health portrait of the customer.
[0017] The updating module is configured to quickly search and acquire real-time health change trend and real-time behavior data of the customer in real time through the LSTM technology, and update the health portrait of the customer regularly.
[0018] The analysis module is configured to acquire a reason for refusal, combine the updated health portrait of the customer and the reason for refusal, and analyze the health risk and potential demand of the customer.
[0019] The recommendation module is configured to combine the updated health portrait of the customer, the reason for refusal, the health risk of the customer and the potential demand based on the dynamic recommendation algorithm to construct a personalized insurance product recommendation scheme.
[0020] In a third aspect, a computer device is provided, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above-mentioned personalized insurance product recommendation method based on the LSTM technology when executing the computer program.
[0021] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the above-mentioned personalized insurance product recommendation method based on the LSTM technology when executed by a processor.
[0022] The scheme realized by the personalized insurance product recommendation method and device based on the LSTM technology, the computer device, and the storage medium can collect original data of a customer, and preprocess the original data, wherein the original data of the customer at least includes health data, living habit data, and historical insurance data of the customer; the original data after preprocessing is subjected to structured processing based on the LSTM technology, and then the original data and the LSTM processing result are stored in a distributed storage system; the health data, the living habit data, and the historical insurance data of the customer in the distributed storage system are extracted, and are integrated into a complete customer data set; based on a machine learning algorithm, data features of the health data, the living habit data, and the historical insurance data of the customer in the customer data set are extracted respectively, and a health portrait of the customer is constructed by training the data features; through the LSTM technology, real-time health change trends and real-time behavior data of the customer are quickly searched and obtained, and the health portrait of the customer is regularly updated; reasons for refusal are collected, and the health risks and potential needs of the customer are analyzed in combination with the updated health portrait of the customer and the reasons for refusal; based on a dynamic recommendation algorithm, an individualized insurance product recommendation scheme is constructed in combination with the updated health portrait of the customer, the reasons for refusal, the health risks, and the potential needs of the customer. In the present application, the scheme takes real-time health data and behavior habits of the customer as the core, and the recommended insurance product always matches the current health status, risk characteristics, and potential needs of the customer (such as recommending a better product after health improvement, and recommending a suitable alternative product after refusal), thereby avoiding “blind recommendation”, allowing the customer to feel targeted service, and improving customer experience and satisfaction; in combination with the dual logic of collaborative filtering (reference to similar customers) and content recommendation (matching own needs), the reference value of group preferences is taken into account, and individual core needs are anchored, thereby reducing invalid recommendations, improving the acceptance of the recommended product by the customer, and improving the accuracy of insurance recommendation and conversion rate; by continuously monitoring health indicators and correlating insurance recommendation adjustment (such as obtaining a better insurance option after health improvement), a positive incentive of “health behavior→health improvement→insurance benefit optimization” is formed, thereby indirectly promoting the customer to actively manage health and enhancing the customer's enthusiasm for health management; based on the recommendation of the accurate health portrait and risk analysis, subsequent claim disputes caused by “mismatch between product and customer risk” can be reduced; at the same time, the dynamic adjustment mechanism can respond to changes in the state of the customer in a timely manner, reduce resource waste caused by “mismatched recommendation”, and reduce the operation and risk costs of the insurance company; the customer can obtain insurance protection that meets the needs of the customer at the moment, thereby avoiding “buying the wrong insurance” and “insufficient insurance”; and the insurance company can more efficiently reach target customers, improve business efficiency, establish a long-term service relationship centered on the customer, and realize two-way adaptation of the customer and the insurance company. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0024] Figure 1 is an application environment schematic diagram of the personalized insurance product recommendation method based on the LSTM technology in an embodiment of the present application;
[0025] Figure 2 is a flow schematic diagram of the personalized insurance product recommendation method based on the LSTM technology in an embodiment of the present application;
[0026] Figure 3 is a structure schematic diagram of the personalized insurance product recommendation device based on the LSTM technology in an embodiment of the present application;
[0027] Figure 4 is a structure schematic diagram of the computer device in an embodiment of the present application;
[0028] Figure 5 is another structure schematic diagram of the computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0030] The personalized insurance product recommendation method based on the LSTM technology provided by the embodiments of the present application can be applied in, for example, Figure 1application environment, wherein the client acquires original data of the customer and pre-processes the original data, wherein the original data of the customer at least includes health data, living habit data and historical insurance data of the customer; the pre-processed original data is subjected to structured processing based on LSTM technology, and the original data and the LSTM processing result are stored in a distributed storage system; the health data, the living habit data and the historical insurance data of the customer in the distributed storage system are extracted and integrated into a complete customer data set; based on a machine learning algorithm, data features of the health data, the living habit data and the historical insurance data of the customer in the customer data set are extracted respectively, and a health portrait of the customer is constructed by training the data features; real-time health change trend and real-time behavior data of the customer are quickly searched and acquired by using the LSTM technology, and the health portrait of the customer is updated regularly; the reasons for refusal are collected, and the health risk and potential demand of the customer are analyzed by combining the updated health portrait of the customer and the reasons for refusal; based on a dynamic recommendation algorithm, a personalized insurance product recommendation scheme is constructed by combining the updated health portrait of the customer, the reasons for refusal, the health risk and the potential demand of the customer. In the present application, the real-time health data and behavior habit of the customer are taken as the core of the scheme, and the recommended insurance product always matches the current health status, risk characteristics and potential demand of the customer (such as recommending a better product after health improvement, and recommending a suitable alternative product after refusal), thereby avoiding “blind recommendation”, allowing the customer to feel targeted service, and improving customer experience and satisfaction; the dual logic of collaborative filtering (reference to similar customers) and content recommendation (matching individual needs) is combined, the reference value of group preference is taken into account, and the core needs of individuals are anchored, thereby reducing invalid recommendations, improving the acceptance of the recommended product by the customer, and improving the accuracy of insurance recommendation and conversion rate; by continuously monitoring health indicators and correlating insurance recommendation adjustment (such as obtaining better insurance options after health improvement), a positive incentive of “health behavior→health improvement→insurance benefit optimization” is formed, which indirectly promotes the customer to actively manage health and enhances the customer's enthusiasm for health management; based on the recommendation of the accurate health portrait and risk analysis, the subsequent claim disputes caused by “mismatch between product and customer risk” can be reduced; at the same time, the dynamic adjustment mechanism can respond to changes in the customer's state in a timely manner, reduce the resource waste caused by “mismatched recommendation”, and reduce the operation and risk cost of the insurance company; the customer can obtain insurance protection that meets his current needs, avoiding “buying the wrong insurance” and “not being fully insured”; the insurance company can more efficiently reach the target customers, improve business efficiency, establish a long-term service relationship centered on customers, and realize the two-way adaptation of customers and the insurance company. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail below through specific embodiments.
[0031] Please refer toFigure 2 As shown, Figure 2 A flowchart of a personalized insurance product recommendation method based on an LSTM technology provided by an embodiment of the present application is shown in FIG. 1. The method includes the following steps:
[0032] S10: Collecting original data of a customer and preprocessing the original data, wherein the original data of the customer at least includes health data, living habit data and historical insurance data of the customer.
[0033] The health data of the customer refers to basic data reflecting the physiological health state of the customer, which is mainly used for assessing health risks and designing health products (such as medical insurance and critical illness insurance). The specific data includes: basic physiological indicators: age, gender, height, weight (and BMI value), blood type, heart rate, blood pressure, blood sugar, blood lipid, etc.; past health records: past medical history (such as chronic diseases such as hypertension, diabetes, and tumors), surgical history, allergy history, family genetic disease history, etc.; recent health status: recent physical examination report (such as liver function, kidney function, and imaging examination results), vaccination record, current medication, etc.
[0034] Collection channels: actively provided by the customer: such as health notification questionnaire filled in when applying for insurance, physical examination report uploaded by the customer, and health information notified during offline consultation; authorized acquisition: after the customer agrees, data is retrieved from medical institutions (hospitals, physical examination centers), health management platforms (such as health APP); device monitoring: real-time collection through cooperative smart devices (such as smart bracelets, blood pressure meters) (which need to be authorized and bound by the customer).
[0035] The living habit data of the customer refers to data related to the daily behavior pattern of the customer, which can indirectly reflect health risks (such as smoking and staying up late, which may increase the probability of getting sick), and is mainly used for personalized service recommendation (such as recommending sports accident insurance for fitness groups).
[0036] Core data content: eating habits: eating structure (such as whether to prefer high-oil and high-salt food), eating frequency (such as whether to have regular three meals), drinking frequency and amount, etc.; work and rest and exercise: sleep duration and quality, exercise frequency (such as the number of times of exercising per week), exercise type (such as slow running, sitting for a long time), commuting method (such as walking, driving), etc.; other habits: smoking status (whether to smoke, smoking years), stress level (such as self-evaluation of work stress), daily travel range (such as whether to often go to high-risk areas), etc.
[0037] Data collection channels: Questionnaire surveys: collected through online questionnaires (such as lifestyle surveys pushed after insurance purchase) or offline interviews; Devices and platforms: obtained from smart devices (such as sleep monitoring bracelets) and lifestyle apps (such as diet tracking apps and fitness apps) (customer authorization required); Behavioral trajectory: after authorization, commuting and travel-related data are obtained from map apps and transportation platforms (such as ride-hailing apps) (anonymized to avoid leakage of location privacy).
[0038] Customer historical insurance data refers to a customer's past behavior records in the insurance field, mainly reflecting their insurance preferences, risk needs, and credit status (such as whether there is any suspicion of insurance fraud). Core data includes: types of insurance previously purchased (such as medical insurance, life insurance, and accident insurance), insured amount, purchase date, and underwriting company; the number of currently valid policies, reasons for surrendered policies, and policy payment records (such as whether they are overdue); claims records: number of claims, reasons for claims (such as illness claims and accident claims), claim amount, and whether there were any disputes regarding the claims; and the completeness of the information provided during the application process (such as whether health information was truthfully filled out), and whether there were any records of being refused coverage or having additional premiums added.
[0039] When the data collection channel is an internal system, it refers to historical insurance data stored in the insurance company's own customer management system (CRM). When the data collection channel is an industry platform, it refers to cross-company data retrieved from insurance industry shared platforms (such as auto insurance information platforms and life insurance claims shared systems) after authorization. When the data collection channel is provided to customers, it refers to other companies' insurance or claims records that customers actively disclose (the authenticity of which needs to be cross-verified with industry platform data).
[0040] The raw data may contain issues such as incompleteness, errors, and disordered formatting (e.g., missing blood pressure values in health data, or duplicated exercise records in lifestyle data). Preprocessing is necessary to transform it into clean, standardized, and usable data for subsequent analysis. The preprocessing steps are as follows: First, data cleaning, primarily addressing errors, missing values, and duplicates. Missing values are typically identified by statistically analyzing the percentage of missing values in each field (e.g., 30% missing for "blood pressure" and 10% missing for "exercise frequency"). Targeted data processing methods include: for critical data (e.g., age and past medical history in health data), contacting the client to supplement the information, or temporarily using industry averages (e.g., average blood pressure for the same age group) (marked as "estimated value"); for non-critical data (e.g., dietary preferences), if the missing percentage is low (<5%), the record can be deleted directly; if the percentage is high, "missing" can be treated as a separate category (e.g., "Dietary preferences: Not filled in").
[0041] Errors in handling error values are identified through logical checks (e.g., "age = 150 years old" is obviously wrong) and format checks (e.g., "blood pressure value" is filled in as "abc" instead of a number). Error values are corrected or removed by contacting the customer to confirm the correction (e.g., "age 150 years old" is actually "50 years old"). Error data that cannot be corrected (e.g., meaningless characters) is directly removed.
[0042] The system handles duplicate values by identifying duplicates by checking for completely duplicate records (e.g., a customer submitting a questionnaire twice) or the same data being collected multiple times (e.g., a smart bracelet repeatedly uploading the same heart rate for the same period). The system also handles deduplication by retaining the latest record (e.g., retaining the last submitted questionnaire) or merging duplicate fields (e.g., taking the average of multiple heart rate readings).
[0043] Data integration primarily involves consolidating data from multiple sources and in various formats. The original data may originate from different channels (e.g., medical examination reports are in PDF format, smart bracelet data is in Excel spreadsheets, and questionnaires are text), and needs to be integrated into a unified format (e.g., a structured data table). Format unification includes: unstructured data conversion (extracting "blood glucose levels" and "liver function results" from PDF medical examination reports into numerical values, and converting "smoking: yes" in text questionnaires into numerical identifiers such as "1" (yes) and "0" (no); and unit unification (e.g., unifying "weight" from "kilograms" and "jin" to "kilograms," and using "blood pressure" as the sole unit).
[0044] Field association primarily uses unique identifiers (such as customer ID) to bind data from different sources (e.g., health data, lifestyle data, and insurance data of "customer ID=123" are merged into the same record); to avoid field conflicts, if the same field (such as "age") is inconsistent in the questionnaire and the ID card, the authoritative source (such as the ID card) shall prevail.
[0045] Data transformation involves converting data into an analyzable format and standardizing it: For continuous data (such as height and weight), normalization (compressing values to the range of 0-1) or standardization (converting to normally distributed data with a mean of 0 and a standard deviation of 1) is performed to avoid the impact of unit differences on analysis (e.g., "height 180cm" and "weight 70kg" cannot be directly compared; standardization allows for a unified dimension); for categorical data (such as "insurance type"), encoding is performed: numbers are used instead of text (e.g., "medical insurance = 1, life insurance = 2") to facilitate subsequent model calculations.
[0046] Construction of derivative indicators: Generate new indicators based on the original data (more valuable for analysis): "BMI index" (to assess obesity risk) is derived from "height and weight"; "health behavior score" (e.g., exercising ≥3 times a week and sleeping ≥7 hours a night is a "high health score") is derived from "exercise frequency and sleep duration"; "insurance stability" is derived from "number of insurance purchases and number of claims" (e.g., ≥2 insurance purchases in the past 3 years is a "stable insurance user").
[0047] Data anonymization and compliance checks primarily ensure data security. Anonymization involves removing or replacing sensitive information, such as replacing "ID number" with "customer ID" and "name" with "anonymous user," retaining only non-privacy fields required for analysis (such as age and insurance records); and blurring information, such as changing "specific address" to "city" and "precise income" to "income range (5k-10k)."
[0048] Compliance review: Check for unauthorized data: Confirm that all data is within the scope of the client's authorization, and remove unauthorized sensitive information (such as unauthorized medical record details); Retain authorization records: Store the client's authorization documents and data together for regulatory verification.
[0049] S20: Based on LSTM technology, the preprocessed raw data is structured and then the raw data and LSTM processing results are stored in a distributed storage system.
[0050] LSTM (Long Short-Term Memory) is a special type of recurrent neural network (RNN) that primarily captures "long-term dependencies" in time-series data (e.g., the correlation between a customer's blood pressure elevation three months ago and current blood sugar fluctuations, or the correlation between a six-month history of staying up late and recent health risks). "Long-term memory" refers to the ability of the network structure (e.g., forget gates, input gates, output gates) to avoid the "short-term memory bottleneck" of traditional RNNs, allowing it to remember early key information in the sequence and use it for subsequent analysis. In customer behavior data scenarios (where preprocessed health data and lifestyle data are mostly time-series data, such as "daily blood pressure values" and "weekly exercise duration"), LSTM's role is to perform "feature extraction" on the time-series preprocessed data (e.g., identifying long-term trends in health data: whether blood pressure continues to rise); to perform "structural transformation" of the data (e.g., converting unstructured time-series records into feature vectors usable for analysis); and to provide "structured intermediate results" for subsequent risk prediction (e.g., disease risk) or behavioral analysis.
[0051] Specifically, in the LSTM processing of customer health data, the long-term trends of health indicators (such as whether blood pressure continues to rise) and the impact of key nodes (such as changes in indicators after a physical examination) are extracted by the processing objective. Input: Preprocessed health data (blood pressure, blood glucose, and physical examination markers sorted by time, converted into a numerical sequence, where the numerical sequence includes at least: systolic blood pressure, diastolic blood pressure, fasting blood glucose, and physical examination markers (1 = yes / 0 = no)). LSTM learning logic: Forget gate: Ignores short-term fluctuations (such as a temporary increase in blood pressure of 1-2 mmHg in a week); Input gate: Reinforces long-term trends (such as blood pressure rising from 125 to 138 mmHg over 3 consecutive months) and key nodes (such as the consistency between physical examination data and daily data); Output gate: Generates structured features.
[0052] The processing results (structured features) include trend feature vectors, [monthly average rate of increase in systolic blood pressure (1.1 mmHg / month), monthly average rate of increase in diastolic blood pressure (0.8 mmHg / month), and correlation between blood glucose and blood pressure (0.85, positive correlation)]; the processing results (structured features) include the risk label "long-term increase in health indicators (medium risk)" (judged based on trend slope); the processing results (structured features) include mutation point markers with no significant mutations (smooth changes in indicators, no sudden increases).
[0053] LSTM processing of customer lifestyle data aims to extract long-term patterns in habits (e.g., whether exercise increases from less to more) and the impact of behavioral changes (e.g., whether exercise improves synchronously after dietary adjustments). Input: Preprocessed lifestyle data (exercise duration, smoking frequency, and dietary tags sorted by time, converted into numerical sequences). LSTM learning logic: Forget gate: Ignores short-term anomalies (e.g., exercise duration temporarily drops to 0 in a week due to a cold); Input gate: Reinforces behavioral change points (e.g., in week 3 of May 2024, the diet changed from "high-oil and salt" to "light," followed by increased exercise and decreased smoking); Output gate: Generates habit association features. Processing results (structured features): Habit trend feature vector: [Monthly average growth rate of exercise duration (0.8 hours / month), monthly average decrease rate of smoking frequency (4 times / month), degree of improvement in exercise after dietary adjustments (0.9, strong correlation)]; Habit stability label: "Positive shift in healthy habits (high stability)" (based on the duration of increased exercise and decreased smoking); Association rule: "Dietary structure adjustment → increased exercise + decreased smoking" (LSTM learns the temporal dependence of the three).
[0054] The LSTM processing of customer historical insurance data aims to extract the stability of insurance behavior (e.g., whether premiums are paid continuously) and changes in demand (e.g., whether the type of insurance has been expanded). Inputs: Preprocessed insurance data (behavior types sorted by time, associated indicators, converted into a numerical sequence: insurance / claim / payment markers (1 = insurance, 2 = claim, 3 = payment), sum assured (normalized), interval time (number of months since the last behavior)). LSTM learning logic: Forget gate: Ignores non-critical behaviors (e.g., fluctuations in the specific amount of a payment); Input gate: Reinforces long-term behavioral patterns (e.g., continuous payments for 5 years without interruption), and behavioral expansion (from type 1 insurance to type 2 insurance); Output gate: Generates insurance behavior features. Processing results (structured features): Insurance stability feature vector: [Continuous payment period (5 years), number of times insurance type was expanded (1 time), duration of insurance after claims (3 years), probability of surrender risk (0.05, extremely low)]; Demand tag: "Increased protection demand (from medical to medical + critical illness)"; Behavioral cycle: "Annual fixed renewal payment (March is the peak payment month)".
[0055] In the preprocessed raw customer data (health, lifestyle habits, historical insurance data), a large amount of data has "time series attributes," which is mainly processed by LSTM technology. For example: daily blood pressure / blood sugar values and monthly physical examination indicators in health data (sequences changing over time); weekly exercise duration and daily sleep quality in lifestyle habit data (sequences recorded in chronological order); and insurance records and claims timelines sorted by time in historical insurance data (e.g., a time series of "medical insurance purchased in 2020 → claims made in 2022 → renewal in 2023"). The core value of this time series data lies not only in the "values at a single point in time," but also in the "patterns of change over time" (e.g., "a continuous decrease in exercise duration for 6 months" may indicate an increased health risk). By learning the patterns in these sequences, LSTM technology can transform "raw time series data" into "structured features containing long-term trends" (e.g., "continuously rising / stable / declining blood pressure in long-term trend labels"), providing a more valuable data carrier for subsequent storage and analysis.
[0056] Distributed storage systems (such as HDFS, Ceph, and GlusterFS) achieve high-capacity, high-availability, and highly scalable storage for massive amounts of data by distributing data across multiple nodes (servers). Storing the "preprocessed raw data" and "LSTM feature data" after LSTM processing into a distributed system requires a storage scheme designed based on the data characteristics.
[0057] The preprocessed raw data is stored in HDFS, for example, with partitioning rules: hierarchical storage by "data type-customer ID-time" for easy querying by customer and time. Path format: / raw data / health data / customer ID=1001 / year=2024 / month=04 (stores customer 1001's health data for April 2024); similarly: / raw data / lifestyle data / customer ID=1001 / year=2024 / quarter=Q2 (stores lifestyle data for Q2 2024); / raw data / insurance data / customer ID=1001 / year=2024 (stores insurance data for 2024). Storage format: stored in Parquet format (column-oriented, high compression ratio), including timestamps and specific indicator values (such as systolic blood pressure, exercise duration). Example: The health data of customer 1001 is stored in HDFS in the file / raw data / health data / customer ID=1001 / year=2024 / month=04 / health_data_202404.parquet, which contains the time-series records of the customer's blood pressure, blood sugar, etc. in April.
[0058] The LSTM processing results are stored in HBase. The table structure is mainly divided into tables based on data type (health feature table, habit feature table, and insurance feature table), with "customer ID" as the row key and "feature type" as the column family. In the health feature table (health_features), the row key is: customer_id = 1001; column family 1 (trend): column name = monthly systolic blood pressure growth rate, value = 1.1 mmHg / month; column name = blood glucose and blood pressure correlation, value = 0.85; column family 2 (label): column name = risk label, value = long-term increase in health indicators (medium risk). In the habit feature table (habit_features), the row key is: customer_id = 1001; column family 1 (trend): column name = monthly average increase in exercise, value = 0.8 hours / month; column name = impact of dietary adjustments, value = 0.9; column family 2 (label): column name = habit label, value = positive change in health habits. Insurance Features Table: Row Key: customer_id = 1001; Column Family 1 (stability): Column Name = Continuous Payment Years, Value = 5; Column Name = Surrender Risk, Value = 0.05; Column Family 2 (demand): Column Name = Demand Tag, Value = Medical + Critical Illness. Index Design: The primary key index is "Customer ID," supporting fast queries by customer; for time-sensitive raw data (such as health data), a "Time Range Index" is added (e.g., querying data from Q2 2024).
[0059] Access scenarios after storage. Scenario 1: Customer risk assessment mainly calls the LSTM processing results in HBase. Customer 1001's health risk label (medium risk) + habit characteristics (positive change in health habits) → assessment conclusion: "Current health indicators have an upward risk, but lifestyle improvements may mitigate the risk. It is recommended to maintain the existing insurance and increase health monitoring."
[0060] Scenario 2: The insurance plan recommendation mainly calls the insurance features (requirements are medical + critical illness) + health features (medium risk) in HBase → recommendation, "On the basis of existing insurance, add hospitalization allowance insurance (adapted to health risk)".
[0061] Scenario 3: Raw data backtracking mainly involves retrieving customer 1001's raw health data from April 2024 from HDFS → verifying the accuracy of LSTM trend features (e.g., confirming whether the monthly blood pressure growth rate calculation is based on real data).
[0062] The core of LSTM-based structured processing is to extract long-term dependent features (trends, correlations, and behavioral patterns) from three types of time-series data (health, lifestyle habits, and insurance), transforming the original time-series sequences into structured feature vectors or labels. Distributed storage achieves efficient data management through HDFS (archiving raw data) and HBase (high-frequency access to structured results). The combination of these two approaches preserves the integrity of the original data (for backtracking and verification) while leveraging LSTM results to support subsequent risk assessments, solution recommendations, and other scenarios. In the insurance industry, this can directly improve the accuracy of customer service (such as customizing insurance based on long-term health trends).
[0063] S30: Extract customer health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset;
[0064] In distributed storage systems, three types of data are typically stored in different components (e.g., HDFS stores raw data, HBase stores structured data) or on different paths. It's essential to first clarify the data storage location, format, and extraction rules. Before extraction, data must be located based on the distributed storage's metadata (e.g., HDFS's NameNode, HBase's metadata tables). Core information includes: a unique customer identifier; all data must be associated with a "customer ID" (e.g., "customer_id = 10001"), which is the foundation for extraction and integration; the data storage path / table name; and the specific time range for extraction (e.g., "data from the last 3 years," "complete data from 2024") to avoid excessive data volume or missing crucial information.
[0065] Different storage components use different extraction methods, requiring adaptation to their interfaces and query syntax to ensure efficient data retrieval: Health data extraction (stored in HDFS): Extraction logic: HDFS stores data in files (e.g., partitioned by "customer ID-year-month"). Files under the corresponding path need to be scanned by customer ID and time range, and the data is then converted into a structured table. Tools and syntax: Use Hive (a data warehouse tool based on HDFS) to create an external table mapping to the HDFS path, and extract data via SQL queries. Extraction results: Structured data (e.g., DataFrame format) containing fields such as customer ID, timestamp, blood pressure, and blood sugar.
[0066] Lifestyle data extraction (stored in HDFS) follows a similar logic to health data extraction, partitioning by "customer ID-time" and focusing on time-series records of lifestyle habits (such as weekly exercise duration and dietary tags). Tools and syntax are also implemented through Hive mapping and querying. The extracted results contain structured data with fields such as customer ID, week, exercise duration, and number of cigarettes smoked.
[0067] Historical insurance data extraction (stored in HBase) involves using HBase, a column-oriented database, to store data in a "row key (customer ID) + family" structure (e.g., the column family insurance_info stores insurance records). Precise queries by row key are required to retrieve data from the corresponding columns. Tools and syntax primarily utilize Phoenix (HBase's SQL interface) or the HBase API for queries. The extracted results contain structured data with fields such as customer ID, activity time, insurance type, and payment status.
[0068] The three types of extracted data need to be temporarily stored in memory or temporary storage (such as Spark in-memory tables or local databases) for easy integration later. When temporarily storing the data, the integrity of the original fields (such as timestamps and original values) must be preserved to avoid information loss.
[0069] The core of the integration is to link three types of data through "customer ID" and align them according to "time dimension" or "analysis requirements" to ultimately form a dataset containing comprehensive customer information. The integration process consists of three steps: "data alignment," "field fusion," and "redundancy handling."
[0070] Data alignment is primarily based on customer ID and time to establish associations. The time granularity of the three types of data may differ (e.g., health data is "daily", lifestyle habits are "weekly", and insurance data is "event-level"). Therefore, a unified association benchmark must be established first: the core association key primarily uses customer_id (customer ID) as the unique association field to ensure all data points to the same customer. If the time alignment rules require integration by "time series" (e.g., analyzing the association between health and lifestyle habits over a specific period): the time granularity is unified to "weekly" (taking the average value of health data within the week, the weekly summary value of lifestyle habits, and marking insurance data by week to indicate whether there has been any activity). If integration requires integration by "overall customer profile" (e.g., generating static customer files), the differences in time granularity are ignored, and the key indicators of the three types of data are summarized (e.g., the latest value for health data, the average level for lifestyle habits, and the current valid insurance policies for insurance data).
[0071] Field fusion involves merging fields from multiple sources, supplementing missing dimensions, and merging fields from the three data types according to "themes" to form a complete field system for the customer dataset. Redundancy handling primarily involves removing duplicate or irrelevant fields. After integration, redundant information may remain (such as duplicate customer IDs or meaningless temporary fields), requiring cleaning: deleting duplicate fields (e.g., if all three data types contain `customer_id`, only one should be retained after integration); removing irrelevant fields, such as storage paths and temporary indexes introduced during extraction (not required by business needs); and handling conflicting fields. If the three data types contain fields with the same name but different meanings (e.g., the "status" field, where health data refers to physical examination status and insurance data refers to policy status), they need to be renamed to distinguish them (e.g., `health_status`, `insurance_status`).
[0072] To ensure the accuracy of the integrated dataset, data validation is required: The integrated dataset must pass "consistency validation," "integrity validation," and "logical validation" to ensure its usability and avoid impacting subsequent analysis due to integration errors. The validated dataset can be output in its final format (e.g., CSV, Parquet, database table) for later applications. When outputting, the dataset's metadata (e.g., data source, integration time, field descriptions) must be clearly stated for user understanding. Consistency Validation: Validation logic: Check the consistency of the association key `customer_id` across the three types of data (e.g., whether health data with "customer ID=10001" is associated with insurance data with "customer ID=10002"); Validation method: Calculate whether the `customer_id` of the three types of data is unique after deduplication (corresponding only to the target customer). If multiple customer IDs exist, the extraction process must be traced back to investigate errors. Integrity Verification: The verification logic mainly checks whether the core fields of three types of data are included to avoid missing key information; Core Field Checklist: Health data should at least include customer_id, time_stamp, and key health indicators (such as blood pressure / blood sugar); Lifestyle data should at least include customer_id, time_week, and core habits (such as exercise / smoking); Insurance data should at least include customer_id, action_type (behavior type), and insurance_type (insurance type); Handling method: If a core field is missing, the corresponding data needs to be extracted again (such as supplementing the missing insurance type field).
[0073] Logical verification: The verification logic mainly checks the logical rationality between data (such as whether the relationship between health and lifestyle habits conforms to common sense); in the example verification rules, if the lifestyle data shows "smoking 20 times a week" and the health data shows "normal lung function indicators" (this may be reasonable in the short term, but should be marked as abnormal in the long term); if the insurance data shows "currently has a critical illness insurance policy" but the health data shows "no past medical history" (this is logically reasonable and does not require processing); the processing method is to mark the logically abnormal data as "suspicious" and pay close attention to it in subsequent analysis (such as manual review).
[0074] The core process of extraction and integration can be summarized as "location and extraction → association and alignment → fusion and verification": Three types of data are accurately extracted using distributed storage query tools, and integrated using "customer ID" and "time" as the core association keys, ultimately forming a dataset containing comprehensive information on health, lifestyle habits, and insurance coverage. This process requires attention to consistent data granularity (to avoid time discrepancies) and logical verification (to ensure data rationality). The integrated dataset can directly support core scenarios such as customer profile construction (e.g., "health-conscious and self-disciplined customers") and risk assessment (e.g., "low-risk customers can be recommended low-premium products"), providing decision-making support for industries such as insurance and health management.
[0075] S40: Based on machine learning algorithms, extract data features from the customer's health data, lifestyle data, and historical insurance data in the customer dataset, and train the data features to build a health profile of the customer.
[0076] Extract "potential features valuable for health profiling" from the original fields (e.g., "long-term upward trend of blood pressure" reflects health risk more accurately than "single blood pressure value"). Select appropriate algorithms based on the characteristics of the three types of data (unsupervised learning to extract potential features, and supervised learning to extract important features).
[0077] Feature extraction of customer health data primarily focuses on "physiological indicators" and "health records." The core objective is to extract "trend features," "abnormal features," and "basic state features." Commonly used algorithms include time series analysis (such as ARIMA), isolated forest (for anomaly detection), and statistical feature extraction. Extraction targets include: basic health status (e.g., whether BMI is above the standard or whether there are chronic diseases); health trends (e.g., whether blood pressure has been rising continuously in the past 6 months or whether blood sugar fluctuations have increased); and health management behaviors (e.g., whether the frequency of physical examinations meets the standard (e.g., ≥1 time per year is considered compliant)).
[0078] Feature extraction from customer lifestyle data: Lifestyle data primarily focuses on "behavioral patterns" (such as exercise, diet, and sleep patterns). The core objective is to extract features related to "behavioral regularity" and "health relevance" (such as "people who exercise regularly have lower health risks"). Commonly used algorithms include clustering (K-Means) and feature importance analysis (random forest). Extraction targets: Habit health level, such as "whether there are high-risk habits such as smoking or staying up late"; habit regularity, such as "whether exercise is performed at least 3 times a week"; the impact of habits on health, such as "the negative correlation between exercise duration and blood pressure" (the more exercise, the lower the blood pressure).
[0079] Feature extraction from customer's historical insurance data: While historical insurance data does not directly reflect health status, features can be extracted through the "correlation between insurance behavior and health" (e.g., healthy individuals are more likely to purchase accident insurance, while high-risk individuals are more likely to purchase medical insurance). Commonly used algorithms include association rule mining (Apriori) and logistic regression (feature importance). Extraction objectives: Health risk perception, such as "whether they actively purchase health insurance" (reflecting their concern for their own health); health-related insurance behavior, such as "number of disease claims" (indirectly reflecting past health problems); insurance stability, such as "whether they continue to renew health insurance" (reflecting their continued need for health protection).
[0080] The features of the three types of data need to be merged into a "health profile feature set," eliminating redundancy and strengthening correlations (e.g., "good exercise habits" and "normal blood pressure" need to be correlated as "positive health features"). The fusion steps are as follows: First, feature screening (removing irrelevant / redundant features), using "feature importance scoring."
[0081] (Such as the Gini coefficient in random forests) Screening: Retain the top 80% of features that have the greatest impact on "health status" (e.g., remove features weakly related to health such as "number of accident insurance purchases"); Remove redundancy: For example, "blood pressure rise trend" and "monthly average rate of blood pressure increase" are highly correlated (correlation coefficient > 0.9), retaining the more intuitive "blood pressure rise trend". Secondly, feature standardization and weighting: Standardization: Unify features of different magnitudes to the 0-1 range (e.g., "exercise duration" and "blood pressure rise rate"); Weighting: Assign weights based on the degree of influence of features on health (e.g., "chronic disease" weight = 0.3, "exercise habits" weight = 0.2, weights can be automatically learned through expert experience or models (e.g., XGBoost). Finally, output a fused feature set, which should include "core health features" (e.g., chronic diseases, blood pressure trends), "behavioral impact features" (e.g., exercise habits, smoking status), and "insurance-related features" (e.g., health insurance purchase tendency).
[0082] A health profile model is constructed based on fused features. The health profile needs to output "interpretable health labels" (such as "health level" and "risk type"). The mapping relationship between features and labels is trained through supervised / semi-supervised models. Profile target definition (label system): First, define the label system of the health profile (which needs to be combined with business needs, such as the insurance industry focusing on "risk level"). An example is as follows: Health level label (target variable): [low risk (0), medium risk (1), high risk (2)]; Sub-profile labels: [health self-discipline type, sub-health improvement type, chronic disease management type, high risk concern type]. Model selection and training (taking "health level" as an example):
[0083]
[0084]
[0085] Model evaluation and optimization: Evaluation metrics: For classification models, use "accuracy and recall" (e.g., recall rate ≥ 0.9 for high-risk customers to avoid missed cases); for scoring models, use "MAE (mean absolute error)" (e.g., the error between predicted score and actual health status ≤ 5 points); Optimization: If the importance of the "smoking status" feature is low (it should actually be high), the data needs to be checked (e.g., whether smoking data collection is incomplete), features should be re-extracted, and the model iterated.
[0086] From model output to intuitive label-based health profiling, after model training, a customer health profile is generated through "feature mapping + label combination," which needs to consider both "quantitative scoring" and "qualitative labels" to facilitate business applications (such as insurance recommendations and health interventions). The profile output format is (quantitative + qualitative).
[0087]
[0088]
[0089] Profile verification (to ensure accuracy) involves cross-validation with the original data: for example, the "healthy and self-disciplined" label should correspond to "≥3 hours of exercise per week + no smoking + normal blood pressure" in the original data; business logic verification: for example, "high-risk customer" should correspond to "chronic disease + smoking + multiple medical insurance claims", which is consistent with common sense.
[0090] The core of building a health profile based on machine learning is "from data to features, and from features to labels": by fusing time series data, clustering, and association rules, a classification / scoring model is trained to ultimately output a profile in the form of "score + label". This profile can directly support business scenarios—such as recommending medical insurance + health management services to "high-risk sub-healthy customers" and recommending low-premium critical illness insurance to "health-conscious and self-disciplined customers", thus achieving "data-driven precision services".
[0091] Step S40 further includes: extracting health features from health data, extracting behavioral features from lifestyle data, and extracting insurance features from historical insurance data; grouping customers based on clustering algorithms and identifying customer groups with similar health and behavioral features; training customer health features, behavioral features, insurance features, and customer groups based on machine learning algorithms to construct customer health profiles; assessing customer health risks and predicting customer health status based on classification algorithms; and assessing customer health risk levels by combining customer health profiles and customer health status.
[0092] Clustering algorithms are used to segment customers into groups with similar health and behavioral characteristics, making customer profiles more targeted. First, the segmentation criteria are determined by selecting health characteristics (such as chronic diseases and blood pressure trends) and behavioral characteristics (such as exercise habits and smoking status) as core indicators for segmentation (these characteristics directly affect health status). Second, data standardization is implemented, unifying the range of features of different magnitudes (e.g., converting "weekly exercise duration" and "blood pressure value" into values between 0 and 1) to prevent a single feature (such as blood pressure value) from dominating the segmentation results due to its high value. Finally, feature dimensions are simplified. If there are many features (e.g., more than 20), dimensionality reduction is used to retain core information (e.g., retaining key features that explain 80% of the data differences) to avoid cluttered segmentation.
[0093] The clustering process (using the commonly used K-Means as an example) determines the number of clusters by using the "elbow method" to determine the optimal number of clusters (K value)—calculating the "sum of squared errors" (the sum of distances between data points and the center of their respective clusters) for different K values, and selecting the K value where the error decreases more slowly (e.g., when K=4, the error decreases significantly more slowly, meaning four clusters). Clustering is performed based on standardized health and behavioral characteristics, grouping customers with similar characteristics into the same group. For example, customers with "no chronic diseases, exercise ≥3 times per week, and do not smoke" are grouped into one category; customers with "hypertension, little exercise, and smoking" are grouped into another. Group characteristic naming analyzes the common characteristics of each group and names them with intuitive labels. For example: Group 1: No chronic diseases, regular exercise, non-smoker → "Healthy and self-disciplined"; Group 2: Hypertension, little exercise, high-oil and high-salt diet → "High-risk for chronic diseases"; Group 3: No chronic diseases, irregular exercise, insufficient sleep → "Sub-healthy and fluctuating"; Group 4: Multiple abnormal indicators, long-term smoking, almost no exercise → "High-risk requiring intervention".
[0094] Health profile construction (integration of features and segmentation results): Combining three types of data features and customer segmentation results, a complete health profile is constructed, including "health status, behavioral patterns, insurance preferences, and risk warnings." The following four types of information are integrated as the basis for profile feature integration: health features (e.g., "no chronic diseases, stable blood pressure"); behavioral features (e.g., "exercises 3 times a week, does not smoke"); insurance features (e.g., "has medical insurance for 3 years, no disease claims"); and segmentation labels (e.g., "healthy and self-disciplined").
[0095] Profile output dimensions (including examples): The profile should consider both "quantitative indicators" and "qualitative labels" for easy and intuitive understanding: Basic health profile: core labels (e.g., "healthy and self-disciplined group, low risk"); health score (maximum score 100, e.g., 85 points, calculated based on feature weighting: no chronic diseases +20 points, regular exercise +15 points, etc.). Behavioral profile:
[0096] Behavioral advantages (e.g., "exercise frequency is better than 80% of peers"); behavioral risks (e.g., "sleep duration has decreased from 7 hours to 5 hours in the past month, requiring attention"). Insurance profile: Current coverage (e.g., "already has 500,000 RMB in medical insurance, no coverage gap"); insurance preferences (e.g., "prioritizes health-related insurance"). Risk insights and recommendations: Potential risks (e.g., "although currently healthy, insufficient sleep may affect immunity"); health recommendations (e.g., "maintain exercise habits and increase sleep time"); insurance recommendations (e.g., "consider supplementing with sports accident insurance to suit daily exercise needs").
[0097] Health risk assessment and prediction (based on classification algorithms) assesses a client's current health risk and predicts future health trends, providing a basis for risk level classification. Current health risk assessment is based on health and behavioral characteristics: Assessment criteria include the number of chronic diseases, the severity of abnormal indicators, and high-risk behaviors (such as smoking and insufficient exercise); Algorithm logic uses classification models (such as random forests and XGBoost) to learn the correlation between "features and risk levels"—for example, clients with "hypertension + smoking" in historical data are often labeled as "medium risk," and the model will associate such feature combinations with medium risk; Output: Current risk level (e.g., "low risk," "medium risk," "high risk").
[0098] Future health status prediction is based on time-series health data and behavioral changes to predict future trends. Prediction criteria include: long-term trends in health indicators (e.g., continuously rising blood pressure) and changes in behavioral habits (e.g., exercise time decreasing from 3 hours to 1 hour per week). The algorithm analyzes patterns of change using time-series prediction methods—for example, predicting blood pressure values for the next 6 months based on blood pressure data from the past 12 months; it adjusts the prediction results by incorporating behavioral changes such as "reduced exercise" (e.g., reduced exercise may accelerate blood pressure increases). Output results include: future health trends (e.g., "10% probability of elevated blood pressure in the next 6 months") and potential health problems (e.g., "If exercise continues to decrease, the risk of abnormal blood sugar may increase from 5% to 15%").
[0099] A comprehensive health risk assessment (integrating profiling and prediction) combines current risks, future trends, and behavioral impacts to determine the final risk level and guide business decisions. The risk level assessment dimensions and weights are based on the following four dimensions, with a total score (0-100 points) calculated according to their weights: Current health status (40%): Number of chronic diseases (e.g., "No chronic diseases = 0 points, 1 chronic disease = 20 points"), Number of abnormal indicators (e.g., "1 abnormality = 10 points"); Behavioral risk (25%): High-risk behaviors (e.g., "Smoking = 20 points, Insufficient exercise = 15 points"); Future trends (20%): Probability of health indicator deterioration (e.g., "Probability of risk in the next 6 months 10% = 10 points"); Coverage (15%): Coverage gap (e.g., "No gap = 0 points, Large gap = 20 points").
[0100] Risk level classification and application: Risk levels are categorized as follows: Low risk (total score ≤ 30 points), Medium risk (31-60 points), and High risk (≥ 61 points). Examples: Low-risk clients (e.g., those with self-disciplined health): Recommend health management services (e.g., discounted physical examinations), and suitable sports accident insurance; Medium-risk clients (e.g., those experiencing sub-health fluctuations): Suggest behavioral improvement (e.g., increased exercise), and recommend medical insurance with health intervention services; High-risk clients (e.g., those at high risk of chronic diseases): Arrange follow-up visits with a health consultant, recommend high-coverage medical insurance with chronic disease management services.
[0101] S50: Through LSTM technology, it can quickly search and obtain real-time records of customers' health change trends and real-time behavior data, and regularly update customers' health profiles.
[0102] LSTM technology excels at efficiently processing real-time time-series data—it can quickly respond to new data while also combining historical data to identify trends, avoiding misjudgments caused by relying solely on data from a single point in time. Its specific functions are reflected in three aspects. First, it enables instant modeling of real-time data. When new health data (such as blood pressure at a specific moment) or behavioral data (such as exercise duration over a certain period) is input, LSTM automatically incorporates it into existing time-series models. For example, if a customer's blood pressure value at 8 AM is input, LSTM will concatenate this value with blood pressure data from the past 24 hours to form a continuous time-series sequence, rather than viewing this single value in isolation. This processing method allows the model to quickly determine whether "the data conforms to a recent trend" (e.g., whether it continues the slight increase in blood pressure over the past 3 days).
[0103] Secondly, it captures dynamic trends in real time. LSTM uses internal gating mechanisms (forget gate, input gate, output gate) to automatically distinguish between "short-term fluctuations" and "long-term trends" when processing real-time data. For example, if a customer's heart rate temporarily increases due to staying up late, LSTM's forget gate will weaken the impact of this short-term fluctuation; however, if the heart rate is 5 beats per minute higher than the average of the previous week for 5 consecutive days, the input gate will amplify this change, identifying it as an "upward trend in heart rate." Simultaneously, for behavioral data, LSTM can capture gradual changes in behavioral habits—for example, as the habit gradually changes from "exercising 3 times a week" to "exercising 1 time a week," the model will update the "weakening of exercise habits" trend label in real time.
[0104] Finally, there's the rapid identification of key changes. LSTM monitors deviations between real-time data and historical trends. When a deviation exceeds a preset threshold (e.g., a sudden spike in blood sugar levels outside the normal range of the past month), it's marked as a "potential anomaly" and prioritized for further analysis. This mechanism ensures a rapid response to sudden changes in a customer's health or behavior (e.g., a sudden increase in blood pressure or a sudden cessation of exercise), providing crucial information for updating the customer profile.
[0105] Based on real-time trend and dynamic data extracted by LSTM, health profiles need to be updated regularly to maintain their timeliness. The update process must balance incorporating real-time changes with ensuring the stability of long-term features to avoid distortion caused by frequent changes. The core logic of the update is to adjust profile dimensions hierarchically based on real-time trends. For dynamic dimensions in the health profile (such as health change trends and recent behavioral patterns), updates are made directly based on the real-time results extracted by LSTM. For example, if LSTM identifies a "decreasing trend in blood pressure over the past week," the original profile's "stable blood pressure" is adjusted to "recently decreasing blood pressure"; if it captures a behavioral change from "occasional smoking to daily smoking," the profile's "low smoking risk" is updated to "increased smoking risk."
[0106] For "long-term characteristic dimensions" (such as basic health status and core behavioral habits), updates need to be made after combining real-time and historical data. For example, if a customer has no history of chronic diseases, and LSTM monitors blood sugar as high for three consecutive months (rather than short-term fluctuations), a long-term characteristic of "prone to high blood sugar" will be added to the profile; if blood sugar fluctuates only for one week, the long-term characteristic will not be adjusted for the time being, but will only be noted in the "recent risk warning".
[0107] The update frequency needs to be set according to the data characteristics and application scenarios. For high-frequency changing data (such as heart rate, daily exercise), the "Real-time Status" module in the profile can be updated daily; for low- to medium-frequency changes (such as blood pressure trends, exercise habits), the trend analysis results of LSTM can be integrated weekly to update the "Trend Features"; for long-term features (such as chronic disease risk), adjustments should be made monthly based on real-time data and regular physical examination results. After each update, the core change points (such as "exercise habits changed from healthy to average" or "blood pressure trend changed from stable to rising") should be marked by comparing with historical profiles to ensure that users can clearly identify the logic of the profile's changes. At the same time, the original data trajectory and trend analysis records processed by LSTM will be retained during the update process for retrospective verification—for example, when "health risk increases" in the profile, it can be traced back to the real-time evidence of "high blood sugar for two consecutive weeks" identified by LSTM, ensuring the interpretability of the update.
[0108] Processing real-time health and behavioral data using LSTM technology essentially leverages its dynamic modeling capabilities for time-series data to capture real-time trends in customer health and behavior. Regularly updating the health profile involves systematically incorporating these dynamic changes into the profile. This mechanism ensures that the profile reflects the customer's latest status (such as recent health fluctuations and changes in behavioral habits) while maintaining its reliability through layered updates and long-term feature stabilization mechanisms. Ultimately, this allows the health profile to guide immediate services (such as providing health advice for sudden increases in blood pressure) and support long-term decision-making (such as adjusting insurance recommendations based on long-term changes in exercise habits).
[0109] S60: Collect reasons for rejection, and combine them with updated customer health profiles and reasons for rejection to analyze customer health risks and potential needs;
[0110] Insurance companies directly assess customer risk based on underwriting rules. Reasons for rejection are collected using standardized methods to ensure complete and clearly categorized information, providing a reliable basis for subsequent analysis. The data collection channels primarily consist of formal records from the underwriting process: first, rejection conclusions from the insurance company's underwriting system (e.g., "Rejection of critical illness insurance due to history of hypertension"), which typically include clearly defined rejection fields (e.g., rejected insurance type, core rejection indicators); second, supplementary explanations from underwriters (e.g., "Although the customer has no definite chronic disease, multiple blood sugar levels in physical examinations over the past year have been high; rejection is based on comprehensive assessment"), which can supplement potential reasons not covered by the system records. From a content classification perspective, reasons for rejection need to be categorized by "risk type," commonly falling into three categories: First, health-related rejections (the most core type), such as "history of hypertension," "abnormal electrocardiogram in the past 6 months," and "BMI ≥ 30 (obesity)"; second, behavior-related rejections, such as "working in a high-risk occupation (e.g., working at heights)" and "a history of frequent smoking (≥ 10 cigarettes per day)"; and third, other rejections (non-health-related), such as "the insured amount far exceeds the income level (moral hazard exists)" and "a history of insurance fraud." The first two categories are directly related to the client's health and are the focus of subsequent analysis. Accuracy is crucial during data collection: avoid vague statements (e.g., a general record of rejection "due to health reasons" should be specified to concrete indicators, such as "rejection due to persistent systolic blood pressure ≥ 160 mmHg"); also, associate the rejected insurance type with the specific risk focus (e.g., "rejection of critical illness insurance" and "rejection of medical insurance" have different risk orientations; the former focuses more on the risk of major illnesses, while the latter focuses on everyday medical risks).
[0111] A health profile encompasses a client's comprehensive health characteristics (such as chronic disease history and blood pressure trends) and behavioral characteristics (such as exercise habits and smoking status), while the reason for rejection is identified as an "explicit risk point" within this profile, according to underwriting rules. The correlation analysis between the two essentially verifies the rationality of the reason for rejection through the profile and uncovers "hidden risks" not covered by the reason for rejection.
[0112] The specific correlation logic is as follows: First, "direct correspondence," meaning that the reason for rejection can be clearly found in the health profile. For example, if the reason for rejection is "history of diabetes," it is necessary to check whether the "chronic disease record" in the health profile includes diabetes and whether the "blood glucose monitoring trend" shows a long-term high level. If the profile clearly records "5-year history of diabetes with poor blood glucose control," then the reason for rejection is consistent with the profile, and the risk point is clear. If the profile does not record a history of diabetes (possibly because the client did not disclose it truthfully), then it needs to be marked as "information deviation," indicating that further verification is required. Second, "extended correlation," meaning that starting from the reason for rejection, related risks can be discovered through the profile. For example, if the reason for rejection is "persistently elevated blood pressure in the past 3 months," combined with the behavioral characteristics in the health profile such as "weekly exercise time decreased from 5 hours to 1 hour" and "increased frequency of high-salt diet," it can be found that "elevated blood pressure" is not an isolated risk, but is related to the behavioral changes of "reduced exercise + high-salt diet." This correlation can explain the cause of the risk and provide direction for subsequent needs assessment. At the same time, attention should be paid to "contradictions between the reason for rejection and the profile." For example, if the reason for rejection is "BMI≥30 (obesity)", but the health profile shows "BMI=28 (overweight)", then the accuracy of the data needs to be verified (such as whether it is the latest BMI value). If it is indeed that the profile has not been updated (the customer has recently gained weight), then it should be suggested that the "basic health characteristics" of the profile should be updated first.
[0113] The reasons for rejection only reflect the "core risks" that underwriters focus on. Combining these with a health profile can pinpoint a more complete set of health risks—including the direct risks that lead to rejection, the derivative risks related to the direct risks, and the potential risks that may worsen in the future.
[0114] First, there's "direct risk confirmation": identifying the specific health issues corresponding to the reason for rejection and assessing the severity of the risk through a profile. For example, if the reason for rejection is "history of coronary heart disease," combined with the characteristics in the health profile such as "early year's echocardiogram showing worsening coronary artery stenosis" and "easily experiencing chest tightness after daily activities," it can be confirmed that "coronary heart disease is in the progression stage, belonging to high risk." If the profile shows "history of coronary heart disease but stable indicators and no obvious symptoms in the past two years," it indicates that the risk is currently controllable, and rejection may be due to the higher risk level of the insurance type (such as critical illness insurance having strict requirements for a history of coronary heart disease). Second, there's "derivative risk mining": based on the direct risk, through behavioral and health trend characteristics in the profile, identifying factors that may exacerbate the risk. For example, if the reason for rejection is "hypertension," and the profile shows "20 years of smoking + less than 1 hour of exercise per week," the derivative risk is "smoking and lack of exercise may accelerate the occurrence of hypertension complications (such as stroke)." If the profile shows "has quit smoking + exercises 3 times a week," the derivative risk is lower, and the risk is more controllable. Finally, there is the "potential risk warning": combining the long-term trends in the profile to judge new risks that may emerge in the future. For example, if the reason for rejection is "high blood sugar (not meeting the criteria for diabetes)," and the profile shows "increased blood sugar fluctuations in the past 6 months + increased frequency of high-sugar diets," then the potential risk is "may develop diabetes in the next 1-2 years." If the profile shows "although blood sugar is high, it has recently shown a downward trend through dietary control," then the potential risk is lower, and the probability of the risk worsening is small.
[0115] After a client is denied insurance coverage, their needs may extend beyond "alternative protection" to include "health improvement (reducing risk to re-insure)" and "supplementing basic coverage." This requires precise identification based on risk assessment and behavioral and underwriting characteristics within the health profile. First, regarding "alternative protection needs": For the risk gap in the denied coverage, recommend alternative insurance policies with more lenient underwriting. For example, if a client is denied critical illness insurance due to "hypertension," but their health profile indicates "no other chronic diseases, and daily medical expenses are mainly outpatient," then "hypertension-specific medical insurance (covering complications) + outpatient insurance" could be recommended to fill the medical coverage gap. Similarly, if a client is denied accident insurance due to "high-risk occupation," but their profile indicates "no health problems," then "occupation-appropriate accident insurance (excluding only high-risk work scenarios)" could be recommended. Second, regarding "health improvement needs": For the risks leading to the denial of coverage, combine the behavioral characteristics in the profile to recommend health services that can reduce those risks. For example, if a client is denied medical insurance due to obesity, but their profile indicates "willingness to exercise but lack of planning," then their need would be "customized exercise plan + nutritionist guidance" (to improve obesity and meet insurance requirements). If a client is denied critical illness insurance due to a smoking history, but their profile indicates "recent intention to quit smoking," then their need would be "smoking cessation intervention services (such as smoking cessation courses + nicotine replacement therapy)." Finally, regarding "basic coverage supplementary needs": if the denial only applies to a specific type of insurance, based on the current coverage status in the profile, recommend basic insurance types that were not denied. For example, if a client is denied critical illness insurance due to diabetes, but their profile indicates "no medical insurance and heavy family responsibilities," then their need would be "inclusive medical insurance (lenient underwriting, covering diabetes medical expenses) + term life insurance (lower health requirements, covering family responsibilities)." Fourthly, regarding "re-insurance preparation needs": for clients who wish to re-insure, based on risk improvement trends, clarify the conditions that need to be met. For example, if a customer is rejected for insurance due to a BMI ≥ 30, but their profile indicates they have "started losing weight and their BMI has dropped to 29 in 3 months," their needs would be "continued weight loss until BMI ≤ 28 (the underwriting standard for most insurance companies) + regular medical check-up records (proving stable weight)," preparing them for re-insurance after 6 months. By collecting the reasons for rejection and linking them to a health profile, we can accurately identify a customer's "direct risks, derivative risks, and potential risks," and uncover potential needs from dimensions such as "protection replacement, health improvement, and re-insurance." The core value of this process lies in transforming "rejection" from a "service endpoint" into a "precise service starting point"—not only providing a suitable solution for the customer but also providing a basis for the insurance company's subsequent risk intervention and customer management.
[0116] S70: Based on dynamic recommendation algorithms, it constructs personalized insurance product recommendation schemes by combining updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
[0117] The recommendation boundaries are defined based on the reasons for rejection and the health profile. The dynamic recommendation process begins by clearly identifying "which products are absolutely unsuitable," avoiding recommending products that the customer has already rejected or that are inherently ineligible for insurance due to their health condition. This is the fundamental bottom line for recommendations. This step requires establishing a "list of prohibited products" by combining the reasons for rejection and the core health characteristics in the health profile. For example, if a customer is rejected for critical illness insurance due to a "history of hypertension," the algorithm will first filter all critical illness insurance policies with strict underwriting requirements for "history of hypertension" (such as products requiring "no history of hypertension"); simultaneously, it will combine this with information on "persistent systolic blood pressure" in the health profile.
[0118] The characteristic of "≥150mmHg" further filters out critical illness insurance policies that, while not explicitly rejecting hypertension, have strict restrictions on blood pressure values (such as requiring systolic blood pressure ≤140mmHg), to avoid repeated rejections after recommendation. For example, if a customer is rejected for medical insurance due to "obesity (BMI ≥30)," the algorithm will filter all medical insurance policies with strict BMI requirements (such as BMI ≤28); simultaneously, considering the trend of "no decrease in BMI in the past 3 months" in the health profile, it will temporarily not recommend medical insurance policies that "require a BMI below 28 to be insured" (to avoid recommending ineffective products), only retaining medical insurance policies with lenient underwriting for obese individuals (such as allowing BMI ≤32, or products that only increase the premium for obese individuals). Through this step, the algorithm can first eliminate products that are "unlikely to be insured," ensuring that the initial scope of recommendations is "compliant and feasible." Within the filtered product range, risk matching is performed based on health risk to identify protection gaps. The algorithm needs to match products that can cover the core risks based on the customer's health risks (direct risks, derivative risks, and potential risks). The algorithm ensures that recommended products "truly address the customer's risk concerns." For example, a customer's health risks might be "hypertension (direct risk) + potential stroke risk (derived risk, due to uncontrolled hypertension)," with a health profile showing "no other chronic diseases, but no regular blood pressure monitoring." The algorithm will prioritize recommending "medical insurance covering hypertension complications (such as stroke)"—these products not only reimburse daily hypertension medication costs but also cover treatment expenses for complications, accurately matching the core risk; while excluding medical insurance that only covers common illnesses (which cannot specifically address hypertension-related risks). Another example: a customer's health risks might be "high blood sugar (direct risk) + future diabetes risk (potential risk)," with the reason for rejection being "rejection of critical illness insurance due to high blood sugar." The algorithm will focus on products that cover "blood sugar-related medical expenses," such as "prediabetes medical insurance" (which reimburses blood sugar monitoring and blood sugar control medication costs), while also including "diabetes prevention services" (such as regular blood sugar management guidance), matching current risks and proactively addressing potential risks.
[0119] The key to risk matching is "not blindly recommending all types of insurance, but focusing on the risks that customers are most likely to face"—for example, for customers whose health risks are mainly "routine minor medical care", outpatient insurance should be recommended instead of critical illness insurance; for customers whose risks are mainly "sudden accidents" (such as healthy people engaged in light physical labor), even if they have been rejected for insurance due to health-related reasons, accident insurance can still be recommended.
[0120] Based on risk matching, the algorithm optimizes the recommendation details based on potential needs to adapt to customer needs. The algorithm needs to combine the customer's potential needs (such as "wanting a low premium", "needing additional health services", "having a reinsurance plan") to adjust the priority and additional content of product recommendations, so that the recommendations are more in line with the customer's actual needs.
[0121] If a customer's potential need is "limited premium budget" and their health risk is "basic medical risk", the algorithm will prioritize recommending products with "low premiums and high deductibles" (such as an annual premium of 300 yuan and a deductible of 10,000 yuan, suitable for dealing with large medical expenses) among the matched medical insurance options, while excluding products with "high premiums but overlapping coverage". If a customer's potential need is "additional health management services" (inferred from "exercise habits" in the health profile), then when recommending medical insurance, products with "medical examination discounts and exercise guidance" will be prioritized to meet their health management needs.
[0122] If a customer's potential need is to "re-insure critical illness insurance in the future" (inferred from the continued focus on health improvement after being rejected for insurance), the algorithm will add a "convertible benefit" product when recommending currently suitable medical insurance. That is, if the current medical insurance is purchased, it can be converted into critical illness insurance without health declaration in the future if the health status improves (such as blood pressure returning to normal). This satisfies the current protection needs and paves the way for future insurance purchases.
[0123] Furthermore, if a customer's potential need is "comprehensive protection" (inferred from "heavy family responsibilities" in the health profile), even if a certain type of insurance is rejected, the algorithm will recommend a "combination plan"—for example, after being rejected for critical illness insurance, a combination of "medical insurance + term life insurance + accident insurance" will be recommended to cover the three core risks of medical treatment, death, and accident, thus ensuring comprehensive protection.
[0124] The core advantage of dynamic recommendation algorithms is their "non-static output"—when a customer's health profile, health risks, and potential needs change, the algorithm dynamically adjusts recommendations based on these changes, ensuring the recommendations always adapt to the latest situation. For example, if a customer is initially denied critical illness insurance due to high blood pressure, a "high blood pressure-specific medical insurance" is recommended. Three months later, if the health profile is updated (by LSTM tracking that blood pressure has dropped to the normal range), the algorithm will detect this change and automatically adjust the recommendation to "critical illness insurance with lenient underwriting (allowing those with controlled hypertension to apply) + the original medical insurance (retaining medical coverage)," because the customer's health risk has decreased, meeting the requirements for critical illness insurance. Similarly, if a customer's initial potential need is "low premiums," basic medical insurance is recommended. Six months later, if the health profile reveals "increased income" and "higher requirements for coverage," the algorithm will automatically prioritize "high-coverage, broad-coverage medical insurance" while retaining the original product as an alternative, adapting to changing needs.
[0125] The adjustment trigger mechanisms fall into two categories: one is "periodic triggering" (such as monthly adjustments based on updated health profile results); the other is "event triggering" (such as immediate updates to recommendations when a customer's health risk decreases, indicators related to the reason for rejection improve, or the customer proactively inquires about new needs). After each adjustment, the algorithm compares the differences between the old and new plans and explains the reasons for the adjustment to the customer (e.g., "Because your blood pressure has stabilized, we have added a critical illness insurance recommendation available for you"), thereby enhancing the credibility of the recommendations.
[0126] Insurance product recommendations based on dynamic recommendation algorithms are essentially a "customer-centric" dynamic adaptation—clearly defining boundaries through reasons for rejection and health profiles, locking in core coverage through health risks, and optimizing details through potential needs, ultimately outputting a solution that is "insurable, covers risks, and meets needs." The core of "dynamic" lies in the fact that the algorithm constantly tracks changes in the customer's health and needs, transforming recommendations from "one-off suggestions" into "a service that continuously optimizes as the customer's status changes," solving current coverage issues while reserving space for future needs.
[0127] Specifically, step S50 includes: determining the recommendation target based on the customer's health risks and reasons for rejection; constructing a recommendation model based on a collaborative filtering algorithm, combining the customer's health profile, reasons for rejection, the customer's health risks, and potential needs; and generating recommendation results based on the recommendation model, wherein the recommendation results include the recommended insurance products, the reasons for recommendation, and the recommendation priority.
[0128] Following the step of constructing a personalized insurance product recommendation plan based on a dynamic recommendation algorithm, combined with updated customer health profiles, reasons for rejection, customer health risks, and potential needs, the method further includes: using LSTM technology to record customer health change trends and behavioral data to construct a customer long-term memory profile; and dynamically adjusting the recommendation strategy based on the customer's long-term memory profile.
[0129] The steps for dynamically adjusting the recommendation strategy based on the customer's long-term memory profile include: regularly monitoring the customer's health indicators and behavioral data to identify changes in health status; dynamically adjusting the recommendation plan if the customer's health status improves over a period of time; and adjusting the recommendation plan based on the new health profile if the customer's behavioral habits change.
[0130] We regularly track clients' health indicators (such as blood pressure, blood sugar, and physical examination results) and behavioral data (such as exercise duration, smoking status, and dietary structure). By comparing historical data, we identify changes in health status—for example, blood pressure stabilizing from high to normal, or exercise increasing from once a week to five times a week. If we detect improvements in a client's health (such as stable control of chronic disease indicators and a reduction in high-risk behaviors), we dynamically adjust insurance recommendations based on this change. For example, if a client was previously restricted from purchasing critical illness insurance due to high blood pressure, we prioritize recommending critical illness insurance with lenient underwriting after health improvement, while retaining the original suitable medical insurance as a supplement. If a client's behavior changes (such as changing from a sedentary lifestyle to regular exercise, or starting to smoke), we first update their health profile (incorporating new behavioral characteristics and related health impacts, such as the downward trend in blood pressure due to increased exercise), and then adjust recommendations based on the new profile. For example, after behavioral improvement, we recommend insurance products with additional health management services; when behavior deteriorates, we focus on recommending medical insurance that covers potential health risks. The entire process, through a closed loop of continuous monitoring, profile updates, and plan adjustments, ensures that recommendations always align with the client's latest health status and behavioral characteristics.
[0131] The steps of constructing a recommendation model based on collaborative filtering algorithm, combined with the customer's health profile, reasons for rejection, health risks, and potential needs, include: using collaborative filtering algorithm to recommend selective insurance products similar to those chosen by similar customers based on their health profile and historical behavioral data; using content recommendation algorithm to recommend demand-based insurance products matching the customer's potential needs based on their health risks and reasons for rejection; and constructing a recommendation model based on the selective and demand-based insurance products.
[0132] The recommendation logic based on collaborative filtering algorithms is as follows: First, find customer groups with similar health profiles (such as health status and behavioral habits) and historical behaviors (such as past insurance preferences) to the target customers. Then, recommend popular insurance products from this group to the target customers. For example, if multiple "healthy, self-disciplined customers without chronic diseases" have chosen sports accident insurance, then recommend this product to customers of the same type, leveraging the selection preferences of similar groups to improve recommendation relevance.
[0133] The recommendation logic based on content recommendation algorithms is as follows: Focusing on the target customer's health risks (such as the risk of potential complications from hypertension) and reasons for rejection (such as being denied critical illness insurance due to hypertension), the algorithm selects products that match the customer's potential needs (such as coverage for complications and lenient underwriting) based on the insurance product's coverage (such as whether it covers hypertension complications) and underwriting conditions (such as whether it accepts applications from people with hypertension). For example, for a customer with hypertension risk who has been denied critical illness insurance, the algorithm recommends medical insurance that covers hypertension and stroke.
[0134] The final recommendation model integrates the results of the two algorithms: it merges the "selective insurance products" obtained by collaborative filtering (based on similar customer selection) and the "demand-based insurance products" obtained by content recommendation (based on matching one's own needs) – prioritizing the retention of products recommended by both algorithms, supplementing products recommended by only one algorithm by ranking them by relevance, and dynamically adjusting the weights based on the customer's real-time health status to form a final recommendation scheme that takes into account both group preferences and individual needs.
[0135] Recommendation models are constructed using either collaborative filtering or content-based recommendation algorithms. Collaborative filtering recommends insurance products similar to those chosen by similar customers based on their health profiles and historical behavioral data. Content-based recommendation recommends insurance products that match the customer's needs based on their health risks and reasons for rejection.
[0136] As can be seen, in the above scheme, intelligent recommendations of personalized insurance products based on the combination of reasons for rejection and customer health data are made by first collecting and preprocessing the customer's raw data, which includes at least the customer's health data, lifestyle data, and historical insurance data; then, structuring the preprocessed raw data using LSTM technology, and storing the raw data and LSTM processing results in a distributed storage system; extracting the customer's health data, lifestyle data, and historical insurance data from the distributed storage system and integrating the data into a complete customer dataset; using machine learning algorithms, extracting data features from the customer's health data, lifestyle data, and historical insurance data in the customer dataset, and training the data features to construct a customer health profile; using LSTM technology, quickly searching and obtaining real-time records of the customer's real-time health change trends and real-time behavioral data, and regularly updating the customer's health profile; collecting reasons for rejection, and combining the updated customer health profile with the reasons for rejection to analyze the customer's health risks and potential needs; and finally, using a dynamic recommendation algorithm, constructing a personalized insurance product recommendation scheme based on the updated customer health profile, reasons for rejection, customer health risks, and potential needs. This invention achieves a solution centered on customers' real-time health data and behavioral habits, ensuring that recommended insurance products are always aligned with their current health status, risk profile, and potential needs (e.g., recommending better products after health improvement, or recommending suitable alternatives after rejection). This avoids "blind recommendations," allowing customers to experience targeted service and improving customer experience and satisfaction. It combines collaborative filtering (referencing similar customer choices) and content recommendation (matching individual needs), balancing the reference value of group preferences with anchoring to core individual needs, reducing ineffective recommendations, increasing customer acceptance of recommended products, and improving the accuracy and conversion rate of insurance recommendations. Furthermore, it continuously monitors health indicators and adjusts insurance recommendations accordingly (e.g., obtaining better products after health improvement). This system (including insurance options) creates a positive incentive cycle of "healthy behavior → health improvement → optimized insurance benefits," indirectly encouraging customers to proactively manage their health and enhancing their enthusiasm for health management. Recommendations based on accurate health profiles and risk analysis can reduce subsequent claims disputes caused by "product-customer risk mismatch." Simultaneously, the dynamic adjustment mechanism can respond promptly to changes in customer status, reducing resource waste caused by "mismatched recommendations" and lowering the operating and risk costs for insurance companies. Customers can obtain insurance coverage that meets their current needs, avoiding "buying the wrong insurance" or "insufficient coverage." Insurance companies, in turn, can reach target customers more efficiently, improving business efficiency while establishing a customer-centric long-term service relationship, achieving a two-way fit between customers and insurance companies.
[0137] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0138] In one embodiment, a personalized insurance product recommendation device based on LSTM technology is provided, which corresponds one-to-one with the personalized insurance product recommendation method based on LSTM technology described in the above embodiments. For example... Figure 3 As shown, this personalized insurance product recommendation device based on LSTM technology includes a data acquisition module 101, a storage module 102, an extraction module 103, a training module 104, an update module 105, an analysis module 106, and a recommendation module 107. Detailed descriptions of each functional module are as follows:
[0139] The data acquisition module 101 is used to collect raw data from customers and preprocess the raw data. The raw data from customers includes at least: customer health data, lifestyle data, and historical insurance data.
[0140] Storage module 102 is used to perform structured processing on the preprocessed raw data based on LSTM technology, and then store the raw data and LSTM processing results in a distributed storage system.
[0141] Extraction module 103 is used to extract customers' health data, lifestyle data and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset.
[0142] Training module 104 is used to extract data features from the customer's health data, lifestyle data and historical insurance data in the customer dataset based on machine learning algorithms, and to train the data features to build a health profile of the customer.
[0143] The update module 105 is used to quickly search and obtain real-time health change trends and real-time behavior data of customers through LSTM technology, and to regularly update the customer's health profile.
[0144] Analysis module 106 is used to collect reasons for rejection of insurance and, in combination with the updated health profile of the customer and reasons for rejection, analyze the customer's health risks and potential needs.
[0145] The recommendation module 107 is used to build personalized insurance product recommendation schemes based on dynamic recommendation algorithms, combined with updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
[0146] In one embodiment, the device further includes:
[0147] The adjustment module uses LSTM technology to record customers' health trends and behavioral data to build long-term memory profiles; and dynamically adjusts recommendation strategies based on these profiles.
[0148] The monitoring module is used to regularly monitor customers' health indicators and behavioral data to identify changes in their health status. If a customer's health status improves over a period of time, the recommended plan is dynamically adjusted. If a customer's behavioral habits change, the recommended plan is adjusted based on the new health profile.
[0149] In one embodiment, the training module 104 is specifically used for:
[0150] Extract health features from health data, behavioral features from lifestyle data, and insurance features from historical insurance data;
[0151] Based on clustering algorithms, customers are grouped and customer groups with similar health and behavioral characteristics are identified;
[0152] Based on machine learning algorithms, we train customers' health characteristics, behavioral characteristics, insurance characteristics, and customer groups to build customer health profiles.
[0153] Based on classification algorithms, assess customers' health risks and predict their health status.
[0154] Assess the customer's health risk level by combining the customer's health profile with the customer's health status.
[0155] In one embodiment, the recommendation module 107 is specifically used for:
[0156] Determine the target of recommendation based on the client's health risks and reasons for rejection;
[0157] Based on the collaborative filtering algorithm, a recommendation model is built by combining the customer's health profile, reasons for rejection, customer's health risks, and potential needs.
[0158] Recommendation results are generated based on the recommendation model, wherein the recommendation results include recommended insurance products, reasons for recommendation, and recommendation priority;
[0159] Based on collaborative filtering algorithms, and according to customers' health profiles and historical behavioral data, we recommend selective insurance products that are similar to those of other customers.
[0160] Based on content recommendation algorithms, insurance products that match the customer's potential needs are recommended according to the customer's health risks and reasons for rejection.
[0161] A recommendation model is constructed based on selective insurance products and demand-based insurance products.
[0162] This invention provides a personalized insurance product recommendation device based on LSTM technology. It collects and preprocesses raw customer data, including at least health data, lifestyle data, and historical insurance data. The preprocessed raw data is then structured using LSTM technology, and the raw data and LSTM processing results are stored in a distributed storage system. The device extracts the customer's health data, lifestyle data, and historical insurance data from the distributed storage system and integrates them into a complete customer dataset. Based on machine learning algorithms, it extracts data features from the customer's health data, lifestyle data, and historical insurance data in the customer dataset and trains these features to construct a customer health profile. Using LSTM technology, it quickly searches and obtains real-time records of the customer's real-time health trends and behaviors, and periodically updates the customer's health profile. It collects reasons for insurance rejection and, combined with the updated customer health profile and reasons for rejection, analyzes the customer's health risks and potential needs. Finally, based on a dynamic recommendation algorithm, it constructs a personalized insurance product recommendation scheme by combining the updated customer health profile, reasons for rejection, customer health risks, and potential needs. This invention achieves a solution centered on customers' real-time health data and behavioral habits, ensuring that recommended insurance products are always aligned with their current health status, risk profile, and potential needs (e.g., recommending better products after health improvement, or recommending suitable alternatives after rejection). This avoids "blind recommendations," allowing customers to experience targeted service and improving customer experience and satisfaction. It combines collaborative filtering (referencing similar customer choices) and content recommendation (matching individual needs), balancing the reference value of group preferences with anchoring to core individual needs, reducing ineffective recommendations, increasing customer acceptance of recommended products, and improving the accuracy and conversion rate of insurance recommendations. Furthermore, it continuously monitors health indicators and adjusts insurance recommendations accordingly (e.g., obtaining better products after health improvement). This system (including insurance options) creates a positive incentive cycle of "healthy behavior → health improvement → optimized insurance benefits," indirectly encouraging customers to proactively manage their health and enhancing their enthusiasm for health management. Recommendations based on accurate health profiles and risk analysis can reduce subsequent claims disputes caused by "product-customer risk mismatch." Simultaneously, the dynamic adjustment mechanism can respond promptly to changes in customer status, reducing resource waste caused by "mismatched recommendations" and lowering the operating and risk costs for insurance companies. Customers can obtain insurance coverage that meets their current needs, avoiding "buying the wrong insurance" or "insufficient coverage." Insurance companies, in turn, can reach target customers more efficiently, improving business efficiency while establishing a customer-centric long-term service relationship, achieving a two-way fit between customers and insurance companies.
[0163] Specific limitations regarding the personalized insurance product recommendation device based on LSTM technology can be found in the limitations of the personalized insurance product recommendation method based on LSTM technology described above, and will not be repeated here. Each module in the aforementioned personalized insurance product recommendation device based on LSTM technology can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0164] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a personalized insurance product recommendation method based on LSTM technology on the server side.
[0165] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of a personalized insurance product recommendation method based on LSTM technology.
[0166] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0167] The preprocessed raw data is structured based on LSTM technology, and then the raw data and LSTM processing results are stored in a distributed storage system.
[0168] Extract customer health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset;
[0169] Based on machine learning algorithms, data features of customers' health data, lifestyle data, and historical insurance data are extracted from the customer dataset, and the data features are trained to construct a health profile of the customer.
[0170] LSTM technology is used to quickly search and obtain real-time health change trends and real-time behavior data of customers, and to regularly update the health profile of customers.
[0171] Collect reasons for rejection and combine them with updated customer health profiles to analyze customer health risks and potential needs;
[0172] Based on dynamic recommendation algorithms, personalized insurance product recommendation schemes are constructed by combining updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
[0173] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0174] The preprocessed raw data is structured based on LSTM technology, and then the raw data and LSTM processing results are stored in a distributed storage system.
[0175] Extract customer health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset;
[0176] Based on machine learning algorithms, data features of customers' health data, lifestyle data, and historical insurance data are extracted from the customer dataset, and the data features are trained to construct a health profile of the customer.
[0177] LSTM technology is used to quickly search and obtain real-time health change trends and real-time behavior data of customers, and to regularly update the health profile of customers.
[0178] Collect reasons for rejection and combine them with updated customer health profiles to analyze customer health risks and potential needs;
[0179] Based on dynamic recommendation algorithms, personalized insurance product recommendation schemes are constructed by combining updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
[0180] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0181] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0183] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A personalized insurance product recommendation method based on LSTM technology, characterized in that, include: Collect raw customer data and preprocess the raw data, wherein the raw customer data includes at least: the customer's health data, lifestyle data and historical insurance data; The preprocessed raw data is structured based on LSTM technology, and then the raw data and LSTM processing results are stored in a distributed storage system. Extract customer health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset; Based on machine learning algorithms, data features of customers' health data, lifestyle data, and historical insurance data are extracted from the customer dataset, and the data features are trained to construct a health profile of the customer. LSTM technology is used to quickly search and obtain real-time health change trends and real-time behavior data of customers, and to regularly update the health profile of customers. Collect reasons for rejection and combine them with updated customer health profiles to analyze customer health risks and potential needs; Based on dynamic recommendation algorithms, personalized insurance product recommendation schemes are constructed by combining updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
2. The personalized insurance product recommendation method based on LSTM technology as described in claim 1, characterized in that, Following the step of constructing a personalized insurance product recommendation plan based on a dynamic recommendation algorithm, combined with updated customer health profiles, reasons for rejection, customer health risks, and potential needs, the method further includes: By using LSTM technology, we can record customers' health trends and behavioral data to build long-term memory profiles for them. The recommendation strategy is dynamically adjusted based on the customer's long-term memory profile.
3. The personalized insurance product recommendation method based on LSTM technology as described in claim 2, characterized in that, The steps of dynamically adjusting the recommendation strategy based on the customer's long-term memory profile include: Regularly monitor customers' health indicators and behavioral data to identify changes in their health status; If the customer's health condition improves over a period of time, the recommended plan will be dynamically adjusted. If a customer's behavior changes, the recommended plan will be adjusted based on the new health profile.
4. The personalized insurance product recommendation method based on LSTM technology as described in claim 1, characterized in that, The steps of extracting data features from the customer dataset, including health data, lifestyle data, and historical insurance data, based on machine learning algorithms, and training these features to construct a customer health profile include: Extract health features from health data, behavioral features from lifestyle data, and insurance features from historical insurance data; Based on clustering algorithms, customers are grouped and customer groups with similar health and behavioral characteristics are identified; Based on machine learning algorithms, we train customers' health characteristics, behavioral characteristics, insurance characteristics, and customer groups to build customer health profiles.
5. The personalized insurance product recommendation method based on LSTM technology as described in claim 4, characterized in that, After the step of training customer health characteristics, behavioral characteristics, insurance characteristics, and customer groups based on machine learning algorithms to construct customer health profiles, the method further includes: Based on classification algorithms, assess customers' health risks and predict their health status. Assess the customer's health risk level by combining the customer's health profile with the customer's health status.
6. The personalized insurance product recommendation method based on LSTM technology as described in claim 1, characterized in that, The steps for constructing personalized insurance product recommendation schemes based on dynamic recommendation algorithms, combined with updated customer health profiles, reasons for rejection, customer health risks, and potential needs, include: Determine the target of recommendation based on the client's health risks and reasons for rejection; Based on the collaborative filtering algorithm, a recommendation model is built by combining the customer's health profile, reasons for rejection, customer's health risks, and potential needs. Recommendation results are generated based on the recommendation model, wherein the recommendation results include recommended insurance products, reasons for recommendation, and recommendation priority.
7. The personalized insurance product recommendation method based on LSTM technology as described in claim 6, characterized in that, The steps for constructing a recommendation model based on collaborative filtering algorithm, combining customer health profiles, reasons for insurance rejection, customer health risks, and potential needs, include: Based on collaborative filtering algorithms, and according to customers' health profiles and historical behavioral data, we recommend selective insurance products that are similar to those of other customers. Based on content recommendation algorithms, insurance products that match the customer's potential needs are recommended according to the customer's health risks and reasons for rejection. A recommendation model is constructed based on selective insurance products and demand-based insurance products.
8. A personalized insurance product recommendation device based on LSTM technology, characterized in that, include: The data collection module is used to collect and preprocess the raw data of customers. The raw data of customers includes at least: customer health data, lifestyle data and historical insurance data. The storage module is used to perform structured processing on the preprocessed raw data based on LSTM technology, and then store the raw data and LSTM processing results in a distributed storage system. The extraction module is used to extract customers' health data, lifestyle data, and historical insurance data from the distributed storage system, and integrate the data into a complete customer dataset. The training module is used to extract data features from the customer's health data, lifestyle data, and historical insurance data in the customer dataset based on machine learning algorithms, and to train the data features to build a health profile of the customer. The update module is used to quickly search and obtain real-time health change trends and real-time behavior data of customers through LSTM technology, and to regularly update the customer's health profile. The analysis module is used to collect reasons for rejection and, combined with the updated customer health profile and reasons for rejection, analyze the customer's health risks and potential needs. The recommendation module is used to build personalized insurance product recommendation schemes based on dynamic recommendation algorithms, combined with updated customer health profiles, reasons for rejection, customer health risks, and potential needs.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the personalized insurance product recommendation method based on LSTM technology as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the personalized insurance product recommendation method based on LSTM technology as described in any one of claims 1 to 7.