Insurance customer loss early warning method and system fusing causal inference large model learning

By using causal analysis and sequential behavior determination, a method for early warning of insurance customer churn is constructed, which solves the problem that causal relationships are not explored in existing methods, achieves high causal degree prediction and personalized intervention, and improves the foresight and scientific nature of insurance customer management.

CN121685168AActive Publication Date: 2026-03-17HANGZHOU SHUO TAI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202610188961.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-03-17
Estimated Expiration
2046-02-10

AI Technical Summary

Technical Problem

Existing insurance customer churn early warning methods mainly rely on correlation analysis, which fails to fully explore the causal relationships between variables. This results in limitations in the model's ability to cope with variable collinearity, confounding factors, and the evaluation of intervention effects. Furthermore, it is difficult to transfer the model to new customer groups or different insurance scenarios, leading to insufficient model generalization ability and making it difficult to support the needs of personalized early warning and intervention decisions.

Method used

By employing causal analysis and sequential behavior determination methods, and constructing customer characteristic data, causal behavior paths, and risk-driven paths, combined with a set of high-risk paths and customer historical behavior trajectories, core influencing nodes are identified, differentiated early warning suggestions are generated, and personalized intervention strategies are generated through rule templates.

Benefits of technology

It achieves high causal prediction and interpretability analysis of customer churn trends, improves the scientific nature of early warning and the pertinence of intervention strategies, and can effectively identify high-risk customers and provide targeted and actionable triggering strategies in complex and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685168A_ABST
    Figure CN121685168A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent research and judgment of insurance data, and particularly relates to an insurance customer loss early warning method and system fusing causal inference large model learning. The method comprises the steps of collecting multi-source data such as customer behavior data, insurance policy information and claim settlement records and performing standardization processing; constructing a causal map based on the structure equation model, and extracting a high causal degree path set; constructing a multi-layer recurrent neural network model for supervised training by taking the path structure as prior; analyzing and optimizing a model structure through path stability; and finally, in combination with the prediction probability and a path backtracking result, outputting a customer risk level and a reversible intervention node. The method realizes causal interpretable prediction and operable path intervention of customer loss, and is suitable for intelligent customer management and personalized operation scenes in the insurance industry.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent research and judgment of insurance data, and particularly relates to an insurance customer loss early warning method and system fusing a causal inference large model learning. BACKGROUND

[0002] Under the background of digital transformation and intelligent service, the insurance industry is facing multiple challenges such as customer demand diversification, service refinement, and customer loyalty maintenance. Customer loss, as an important indicator to measure the stability and service quality of an enterprise, is paid more and more attention by insurance companies. Especially in the field of long-term products such as life insurance and health insurance, the continuous participation of customers has a profound impact on the income structure and risk assessment system of the enterprise. Therefore, how to use advanced data analysis and artificial intelligence technology to realize the early identification and effective intervention of customer loss has become one of the important directions of insurance technology innovation.

[0003] The existing insurance customer loss early warning method mainly relies on statistical modeling and machine learning technology to build a scoring model, and predicts the loss risk through variables such as historical insurance data, customer behavior characteristics, and policy compliance. Among them, models such as logistic regression, decision tree, random forest, and XGBoost are widely used, combined with customer segmentation strategies such as RFM model and life cycle model to build a loss scoring system. However, these methods are mostly based on correlation analysis and fail to fully explore the causal relationship between variables, which has certain limitations in dealing with variable collinearity, confounding factor interference, and intervention effect evaluation. In addition, existing models often rely on a large amount of labeled data for supervised training, requiring high data quality, sample balance, and feature engineering level, and are difficult to migrate to new customer groups or different insurance scenarios, resulting in insufficient model generalization ability and difficulty in supporting personalized early warning and intervention decision-making needs.

[0004] Under the dual demands of insurance refinement operation and intelligent decision-making, it is urgent to introduce a technology path with better causal explanation ability and generalization ability to build a customer loss early warning method applicable to complex dynamic environments, so as to improve the forward-looking nature of insurance customer management and the scientific nature of intervention strategies. SUMMARY

[0005] To solve the above problems, the purpose of the present application is to provide an insurance customer loss early warning method fusing causal analysis and sequential behavior determination, comprising the following steps: S1, collection of customer feature data and construction of behavior indicators: collect multi-source customer behavior data including login behavior, policy renewal behavior, claim records, payment operations, online interactions, etc., form a structured customer feature data set after cleaning, and construct a behavior indicator subset including interaction frequency, stay duration, operation activity, etc. S2, construction of causal behavior path and screening of risk-driven path: based on customer behavior index data, analyze the key variable causal relationship affecting the customer policy state change, construct several causal path diagrams, test the behavior disturbance of each path, and screen the paths with obvious differences before and after the behavior state change as the high-risk path set; S3, customer loss risk state reasoning and behavior trend determination: combining the high-risk path set and the customer historical behavior trajectory, the loss trend of each customer is inferred, the core influence nodes in the path are identified, and the future one-period policy state risk level of the customer is determined according to the behavior evolution direction of the customer in the renewal period; S4, early warning level classification and behavior intervention point calibration: the customer loss trend evaluation results are divided into non-risk, medium-risk and high-risk levels, and the key behavior nodes and turning variables of the customers in the risk state are identified, and are calibrated as "intervenable path nodes" or "irreversible behavior nodes", and differential early warning suggestions are generated based on the rule templates; S5, risk customer output and artificial intervention suggestion pushing: the identified risk level customers are grouped and managed, the risk level, influence path and node description of each customer are output, and intervention suggestion documents for customer service or channel are generated, including renewal communication prompts, personalized marketing reminders or claim handling priority suggestions, etc.

[0006] As a preferred technical solution, the construction of the behavior index in S1 includes: S11, extracting the customer access log and click flow record of the past three months according to the time stamp, and constructing the user behavior sequence; S12, standardizing the behavior sequence according to the time span and behavior distribution law to form a balanced time step data set; S13, performing interval judgment and filling on the behavior abnormal segment, and using the completed data as the main input vector and the original abnormal interval to construct an additional score factor; S14, scoring the behavior activity stability through the scoring function to determine whether the current behavior state of the customer has faults or fluctuations.

[0007] As a preferred technical solution, the construction of the causal path in S2 includes: Taking the insurance customer behavior variables as nodes, the causal edges are selected according to the structural correlation degree to form a directed behavior path diagram; The behavior disturbance test method is used to change the behavior frequency of the variables in the path, and the change of the loss probability distribution before and after the change is calculated; The change of the loss probability caused by the behavior disturbance is used as the path effectiveness criterion, and the high-sensitivity path is selected as the core early warning path set.

[0008] As a preferred technical solution, the determination of the customer churn risk in S3 comprises: Based on the current node state of the customer in the high causality path, combined with the historical evolution trend of the path, the possibility of developing to a non-renewal state is predicted; If the similarity between the current behavior of the customer and the typical churn customer behavior in the early warning path is higher than a set threshold, it is determined that there is a potential churn risk.

[0009] As a preferred technical solution, the risk level determination rule in S4 comprises: If the predicted churn probability is greater than a preset second threshold, the customer is marked as a high-risk customer; If the prediction result is between the first threshold and the second threshold, the customer is marked as a medium-risk customer; Otherwise, the customer is marked as a non-risk customer; Wherein, the threshold can be automatically calibrated based on historical data or set by operation management rules.

[0010] As a preferred technical solution, the calibration of the behavior intervention point in S4 comprises: According to the comparison between the current behavior node of the customer and the early warning path structure, the path node with the possibility of behavior reversal is identified, and the node is calibrated as an intervenable node; If a behavior sequence is interrupted, the customer has no active signal or the access path is completely consistent with the risk customer, then the node is calibrated as an irreversible node; The two types of nodes are output respectively for personalized intervention recommendation and risk closed-loop management.

[0011] As a preferred technical solution, the artificial intervention suggestion pushing in S5 comprises: Generate a customer-level risk profile, including customer identification, behavior label, risk level and core path node information; According to the customer's historical renewal state and current behavior state, generate corresponding intervention suggestion scripts for each type of customer; Synchronize the high-risk customer list to the sales, customer service, and agent channel systems to trigger exclusive communication or task reminder processes.

[0012] The application also provides an insurance customer churn early warning system integrating a causal inference learning mechanism, for implementing the method, characterized by comprising: A customer data acquisition module for synchronously acquiring customer behavior data from a customer management system CRM, a policy management system, an online platform, and the like; A causal path analysis module for establishing a path structure diagram based on behavior variables and calculating a causal score; A risk state reasoning module for predicting a future renewal state of a customer and classifying a risk level; Node identification and intervention suggestion module: for generating key behavior nodes and corresponding early warning suggestion information; Risk output and task linkage module: export high-risk customer data to external intervention system or manual intervention platform.

[0013] Beneficial effects: The application realizes the causal structure identification and reasoning path control which is difficult to obtain in traditional churn prediction by introducing structural equation modeling and path scoring mechanism, can structurally model the potential causal relationship between high-dimensional heterogeneous customer characteristics, and clearly define the conduction logic chain between behavior variables. Compared with the shallow learning model based on correlation only, this method can not only be used for predicting whether the customer churns, but also can output the quantitative explanation of "why churns" based on the graph, so that the prediction result has strong traceability, and provides clear path basis for subsequent business intervention.

[0014] The deep reasoning model of the application fuses high-causality path and multi-layer recurrent neural network, has double structural characteristics: on the one hand, the causal path is reserved as semantic prior, which is used to limit the structural boundary of information propagation; on the other hand, the behavior transition in time evolution is dynamically modeled through recursive learning. The structure significantly improves the model's ability to perceive behavior disturbance and capture long-term fluctuation trend, realizes model stability control based on path fluctuation analysis, and thus guarantees the stable reasoning performance of the system in non-static scenes such as business cycle fluctuation and user behavior evolution.

[0015] The application establishes a linkage closed loop from prediction to intervention through the identification mechanism of "reversible causal disturbance node", so that the system can not only judge whether the customer has churn risk, but also clearly point out which behavior nodes are still in the adjustable state, and output the variable set with the highest intervention potential through the path backtracking mechanism. The mechanism provides a directional and operable trigger strategy for the customer operation team, realizes the paradigm shift from "static classification" to "dynamic reversal" of churn risk management, and has high practical value and system integration advantage in the insurance customer life cycle operation scene. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 The method flowchart of the application. DETAILED DESCRIPTION

[0017] In order to deepen the understanding of the application, the application will be further described in combination with the embodiments below, and the embodiments are only used to explain the application and do not constitute a limitation on the protection scope of the application.

[0018] Embodiment one As Figure 1As shown, this embodiment proposes an insurance customer churn early warning method that integrates causal inference large model learning. Based on multi-source heterogeneous customer data, structured causal inference graph and deep recursive learning model, an intelligent insurance decision support process oriented towards prediction, analysis and intervention is established.

[0019] The method described is applicable to the integrated deployment of customer service systems, marketing strategy systems, and risk control systems in the insurance industry, and is scalable and platform compatible.

[0020] This method includes the following main steps: S1. Collection and Standardization Processing of Customer Feature Data: This step aims to construct a customer feature dataset with a unified structure and consistent dimensions, serving as the input foundation for the model. The data acquisition module integrates access to multiple business systems, and the main sources of the collected data are as follows: Customer behavior data is derived from client interaction logs, including access records from PCs and mobile devices, clickstream logs, page navigation behavior, dwell time, mouse paths, and form completion information. These logs are recorded by the Web Behavior Analytics SDK and uploaded to a log database for daily batch processing.

[0021] Historical policy information comes from the core policy system, including fields such as each customer's policy number, product type, coverage amount, payment frequency, policy start and end dates, whether it is a renewal, and cumulative payment duration.

[0022] Claim records are sourced from the claims subsystem, including application time, payout amount, payout status (under review, rejected, closed), number of claims, classification of reasons for compensation, and whether it is a critical illness.

[0023] Online interaction data comes from the customer service center and user feedback system, including text chat logs, telephone service tags, satisfaction ratings, complaint types, and customer-initiated feedback.

[0024] Marketing response behavior data comes from the marketing campaign system: such as whether customers clicked on SMS links, opened marketing emails, participated in promotional activities, or claimed coupons.

[0025] After data collection is complete, follow the standardized procedure below: Time alignment processing: unifies the features of different time zones and time granularities into a time series in days; Missing value handling: For time series data, a sliding window average interpolation is used; for non-time data, the mean of the same type of customer is used for filling. Feature normalization: Min-Max normalization is used to compress numerical features into 0-1 values, and one-hot encoding is used for categorical variables. Behavioral segmentation: Customer behavior throughout the entire lifecycle is divided into three levels: the past 7 days, the past 30 days, and the past 90 days, and high-frequency, medium-term, and long-term behavioral trend features are extracted respectively. Anomaly detection: An algorithm based on the Local Outlier Factor (LOF) is introduced to perform quality verification on the feature data and remove obviously erroneous records.

[0026] Ultimately, a unified set of customer feature vectors is obtained, which serves as the basic input for subsequent causal modeling and prediction models.

[0027] S2. Causal Graph Construction and High Causality Path Selection: After standardizing customer feature data, the process moves on to constructing a causal graph. This is the first core innovation of this method, aiming to extract paths with clear causal relationships from a large number of feature variables and use them as prior knowledge for subsequent inference models.

[0028] Structural Equation Modeling (SEM): This step uses structural equation modeling as the basis for causal inference, constructing a linear equation of the following form for each pair of variables:

[0029] Where y is the dependent variable, x is the independent variable, β is the dependency coefficient, and ε is the error term. The parameter matrix of the model is solved using the maximum likelihood estimation method.

[0030] Construct a causal graph structure: Treat all variables as nodes in the graph, and retain edges with dependency coefficients greater than a threshold (such as 0.3) to form a preliminary directed graph. The edges in the graph represent the causal directionality between variables.

[0031] Calculate path score: Construct a multi-hop path from the cause variable to the target variable (e.g., "whether to renew insurance"), and calculate the path score for each path. Dependency score (weighted average of edge coefficients in the path); Relevance score (cooperation variance between path nodes); Path entropy (reflects the amount of information in a path); Path compression (the ratio of non-redundant nodes in the path).

[0032] Path filtering and pruning: Paths with scores below a set threshold (e.g., path entropy < 0.5, dependency score < 0.4) are deleted, while paths with high structural redundancy are merged.

[0033] Causal intervention simulation: Bayesian networks are used to conduct intervention tests on each path, simulating the distribution of customer behavior before and after the change of path node state, and KL divergence is used to evaluate the degree of impact of path intervention.

[0034] Finally, the set of paths that rank in the top 10% and have significant intervention effects (KL divergence) are selected as the "high causal path set" and input into the next step of the modeling process.

[0035] S3. Recursive Structure Learning and Customer Intent Reasoning Modeling: This step, based on an RNN-like deep learning structure, jointly models customer behavior time series and high-causality path sets to achieve recursive reasoning of customer behavior states towards future churn probabilities.

[0036] Input data construction: customer behavior time series tensor, path structure matrix, label vector; Model architecture design: The first layer is the path structure encoding layer, which uses a graph embedding algorithm to map the path matrix into a vector representation; The second layer is the BiGRU time series processing layer, which performs forward and backward analysis on the customer behavior sequence. The third layer is the fusion layer, which integrates the path vector and the behavior vector with attention weighting. The fourth layer is a fully connected sigmoid output layer that outputs the dropout probability.

[0037] Training details: Loss function: Binary cross-entropy is used; Optimizer: Adam, initial learning rate 0.001; Training rounds: 30 rounds, with each round having a Mini-Batch size of 64; Overfitting prevention strategy: Dropout=0.2, EarlyStopping tolerance step count is 5; Model evaluation metrics: AUC, F1-Score, Precision@Top5.

[0038] This model can not only output the churn probability of each customer, but also trace the high causal path on which its prediction results depend, thus possessing good interpretability.

[0039] S4. Path stability analysis and model optimization: This step aims to improve the inference model's adaptability to customer behavior drift by identifying unstable paths and performing targeted training optimization through multi-time period validation and path disturbance response analysis.

[0040] Customer grouping: The sample customers are divided into multiple 30-day sliding windows according to time, and each group of data is used as a stability assessment unit; In-window inference simulation: Perform a prediction independently within each time window and record the average participation frequency, prediction accuracy, and output probability volatility for each path; Constructing a fluctuation score matrix: Construct a matrix for the performance of each path under different time windows, and calculate its mean square deviation as the path fluctuation score; Screening unstable paths: If the mean squared deviation of a path exceeds twice the mean score of the overall path, the path is marked as an unstable path and added to the training reinforcement list. Path weight reallocation: The sampling weights of each path sample in the training dataset are readjusted, maintaining the default sampling probability of 1.0 for stable paths and increasing the sampling weight of unstable paths by 1.5 times; Optimize model retraining: Re-execute 10 rounds of iterative training to improve the overall robustness of the model under the optimized path sample distribution.

[0041] The final result is a set of deep inference models with higher stability, which have time transfer capabilities and cross-cycle generalization capabilities.

[0042] S5. Customer Risk Level Assessment and Reversible Causal Path Output: After the model is deployed, the system can make real-time predictions for any one or more target customers and return the risk level and reversible intervention points. The specific process is as follows: Predictive Invocation: Input the target customer's behavioral data and static attributes from the past 90 days into the model inference interface to obtain the output probability P; Risk classification assessment: If P ≥ 0.75, the customer is marked as "high risk"; If 0.5 ≤ P < 0.75, then it is considered "medium risk"; If P < 0.5, it is considered "low risk".

[0043] Path contribution tracking: By backtracking the attention mechanism and path weights in the model, the path structure that contributes most to the current prediction (Top 3 paths) is determined. Determining the reversibility of causal nodes: Check if there are any variable nodes in the path that can be affected by marketing, service, policy changes, etc. (such as "number of customer service interactions in the past 30 days", "frequency of coupon usage", etc.). If these nodes have experienced significant fluctuations in the past and are currently in an "adjustable state," they are marked as "reversible intervention nodes."

[0044] Intervention suggestion generation: Combining path analysis and node behavior values, specific intervention suggestions are automatically generated, such as "proactively initiate policy renewal communication", "send customized discount reminders", and "push successful claims cases" to improve customer retention probability.

[0045] Output structure: Customer ID; Current risk level; Predicted probability value; Main causal path name (variable sequence); List of reversible intervention nodes; Intervention strategy suggestion text.

[0046] The above outputs can be automatically synchronized to the CRM system, allowing customer operations specialists or marketing teams to implement differentiated relationship maintenance.

[0047] Example 2 This embodiment provides an insurance customer churn early warning device that integrates causal inference large model learning. It is applicable to business scenarios such as insurance customer lifecycle management, renewal strategy formulation, and intelligent customer service operation. It is particularly suitable for deployment in the intelligent operation platform, data middle platform, or private cloud environment of insurance companies.

[0048] This device is designed for real-time processing of large-scale customer behavior data, causal structure learning, and high-dimensional reasoning tasks. It combines causal graph modeling and deep learning models, which can effectively improve the accuracy of customer churn prediction and the targeting of intervention strategies.

[0049] The device includes the following core modules: 1. Customer Data Acquisition and Preprocessing Module: This module is used to collect multi-source customer data and perform format unification, missing data filling, and feature standardization processing. It mainly includes the following sub-components: Data access submodule: It uses Kafka data bus to connect with the policy data table, claims database, interaction log system and marketing management platform of the insurance business system; it supports structured data (SQL tables), semi-structured data (JSON, XML) and unstructured data (customer service voice text) input; it is compatible with scheduled batch retrieval and real-time streaming synchronization, and the data synchronization cycle is configured to be 10 minutes.

[0050] Feature standardization submodule: performs unified field mapping on data from different sources, including field naming standardization, unit conversion (e.g., converting ten thousand yuan to yuan), and type forced conversion; uses the Pandas data framework to complete segmented normalization processing, such as Z-score standardization of click count, login frequency, etc. according to behavioral distribution; null value handling strategies include: forward filling of the most recent time point, filling with the mean of similar customers, and interpolation completion (e.g., behavioral time series).

[0051] Customer behavior vector generation submodule: Constructs fixed-length feature vectors (e.g., 300-dimensional) containing behavioral feature summaries for the last 7, 30, and 90 days; supports generating window sliding behavior fragments on a user-by-user basis, which can then be input into subsequent recursive models.

[0052] This module outputs a standardized customer feature dataset, providing a unified data foundation for the causal graph construction module.

[0053] 2. Causal Graph Construction and Path Selection Module: This module is used to build structural equation models based on customer characteristic data, form a causal relationship graph between variables, and extract a set of high causal paths, mainly including: The causal relationship identification submodule adopts structural equation modeling (SEM) as the causal learning framework; it sequentially regresses all latent independent variables for each target variable (such as "whether it is lost"), calculates the path dependence coefficient and the fitting residual; and uses maximum likelihood estimation to solve for the strength of directed edges between nodes.

[0054] The causal graph structure generation submodule treats all variables as nodes and retains edges with causal strength greater than a set threshold (e.g., 0.3) to form a preliminary causal network; the NetworkX graphics computing library is used to store the directed graph structure.

[0055] The path search and combination submodule: Based on the DFS algorithm, it searches for all causal paths from "key upstream variables" to "whether customers are churned"; for each path, it calculates indicators such as path length, cumulative causal strength, and node information entropy.

[0056] The path scoring and screening submodule performs a simulated intervention experiment for each path and calculates the distribution of customer states before and after the intervention based on Bayesian networks. KL divergence is used as the basis for path scoring, with higher scores indicating a greater causal intervention impact. Paths with scores in the top 10% and path difference not equal to 0 are selected to form a set of high causal paths.

[0057] This module outputs a path structure matrix and a score vector, which serve as structural priors for model training input.

[0058] 3. Inference Model Construction and Training Module: This module is used to construct a recurrent neural network structure based on high-causal paths as priors, and to perform time-series learning of customer status and churn intention prediction, including: The path embedding encoding submodule represents each causal path as a sequence of nodes and uses a Node2Vec or graph neural network encoder to embed the path structure into a low-dimensional dense vector (e.g., 64-dimensional); all paths constitute a path vector set.

[0059] The behavior sequence recursive processing submodule: Customer behavior vectors are arranged into a sequence according to time windows and input into the BiGRU structure; the hidden state sequence H of each time step is output, representing the potential state evolution of the customer at each time node.

[0060] The path behavior fusion and attention mechanism submodule utilizes a weighted attention mechanism between the path vector and the sequence state to calculate the interpretive contribution of the path to the current behavior state; the fused state representation is then input to the fully connected layer.

[0061] The classification output and optimization submodule outputs the probability value of whether a customer will churn in the next cycle (e.g., 30 days) through the Sigmoid activation function; it uses cross-entropy as the loss function and adopts the Adam optimizer for training; it records AUC, F1, and accuracy in each training round for model tuning and evaluation.

[0062] This module supports model saving, loading, fine-tuning, and batch training, and can be deployed independently on GPU-accelerated nodes or AI model services.

[0063] 4. Path stability analysis and model iteration module: This module analyzes the contribution of paths in the model to the stability of the results, eliminates highly volatile paths, and enhances the model's generalization ability. Its structure includes: Time window grouping submodule: Divides the training data into multiple time periods according to the customer's active time, with a typical window length of 30 days; each window serves as an independent evaluation unit.

[0064] Model Inference Simulation Submodule: Runs model inference independently in each window, records the deviation between predicted values ​​and actual results, and marks the participation of each path in inference.

[0065] The path fluctuation scoring submodule: statistically analyzes the performance of each path in inference across all time windows; calculates its standard deviation, variance, and mean error, and constructs the path fluctuation matrix.

[0066] The path weight adjustment submodule sets high training weights for paths with fluctuating scores significantly higher than the average to enhance the model's adaptability to them; and sets low sampling rates for paths with stable scores but low information contribution to reduce interference.

[0067] Retraining control submodule: Retrains the model based on the adjusted weights; supports two strategies: local path freezing or full path retraining.

[0068] This module outputs a set of updated and optimized model structures, further improving prediction accuracy and path reliability.

[0069] 5. Risk Prediction and Reversible Path Output Module: This module is business-oriented, providing a real-time prediction interface and interpretable path output, facilitating manual or automated execution of customer retention strategies. It includes: Customer input receiving submodule: Receives customer identifiers and behavioral data from the past 90 days transmitted from the front end; parses them into a unified format input vector.

[0070] Risk level determination submodule: Based on the probability value output by the model, two thresholds are set (such as 0.5 and 0.75): greater than 0.75 is high risk; between the two is medium risk; less than 0.5 is low risk.

[0071] The path interpretation and reversible node identification submodule: backtracks the attention mechanism results and extracts the Top-N key causal paths; marks the node variables in the path that can be manually or systematically intervened, such as interaction frequency and claim status; determines whether these nodes have historical evidence of changing their status, and if so, marks them as reversible perturbation nodes.

[0072] The intervention suggestion generation submodule generates text-based operational suggestions based on model analysis results and node values, such as: "It is recommended to stimulate the willingness to renew the policy through coupons"; "It is recommended to arrange customer service representatives to proactively contact and communicate about claims disputes"; "It is recommended to add the customer to the short-term recall marketing list".

[0073] The Results Display and Output Interface submodule supports front-end display of risk scoring cards, path maps, and intervention suggestions; and outputs results to customer relationship management systems (CRM) or business operation dashboards via REST API.

[0074] This device achieves a closed-loop process from data acquisition, causal relationship modeling, deep reasoning to risk output and path interpretation through multi-module collaboration. Compared to traditional prediction methods based on shallow models, this device has the following advantages: Structural interpretability: Introducing a causal graph structure improves model transparency; Results are traceable: outputs critical path and node variables to assist business intervention; High performance: It adopts BiGRU and attention mechanism to achieve compressed learning of high-dimensional sequences; Flexible deployment: Supports deployment on private clouds, local servers, and model service platforms.

[0075] This embodiment can be widely applied to insurance technology scenarios such as customer retention analysis, personalized operations, risk stratification, and product pricing.

[0076] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. An insurance customer churn early warning method fusing causal analysis and sequential behavior determination, characterized in that, The method comprises the following steps: S1, collection of customer feature data and construction of behavior indicators: Collect multi-source customer behavior data including login behavior, policy renewal behavior, claim record, payment operation, online interaction, form a structured customer feature data set after cleaning, and construct a behavior indicator subset including interaction frequency, stay time, operation activity; S2, construction of causal behavior path and risk driving path screening: Based on customer behavior indicator data, analyze the key variable causal relationship affecting the change of customer policy state, construct several causal path diagrams, test the behavior disturbance of each path, and screen the paths with obvious differences before and after the behavior state change as the high-risk path set; S3, customer loss risk state reasoning and behavior trend determination: Combine the high-risk path set and the customer historical behavior trajectory to reason the loss trend of each customer's current behavior state, identify the core influence nodes in the path, and judge the future one-period policy state risk level of the customer according to the behavior evolution direction of the customer in the renewal period; S4, early warning level classification and behavior intervention point calibration: Divide the customer loss trend evaluation results into non-risk, medium-risk and high-risk levels, identify the key behavior nodes and turning variables of the customers in the risk state, and calibrate them as "intervenable path nodes" or "irreversible behavior nodes", and generate differentiated early warning suggestions based on the rule template; S5, risk customer output and artificial intervention suggestion pushing: Group the identified risk level customers, output the risk level, influence path and node description of each customer, and generate intervention suggestion documents for customer service or channel, including renewal communication prompt, personalized marketing reminder or claim handling priority suggestion. 2.The insurance customer churn early warning method of fusing causal inference learning mechanism according to claim 1, characterized in that, The construction of behavior indicators in S1 includes: S11, extract customer access logs and click stream records in the past three months according to time stamp, and construct user behavior sequence; S12, standardize the behavior sequence according to time span and behavior distribution law to form a balanced time step data set; S13, perform interval judgment and filling on the behavior abnormal segment, and use the completed data as the main input vector and the original abnormal interval to construct additional scoring factors; S14, score the behavior activity stability through the scoring function to determine whether there is a fault or fluctuation in the current behavior state of the customer. 3.The insurance customer churn early warning method of fusing causal inference learning mechanism according to claim 1, characterized in that, The construction of causal path in S2 includes: Take the insurance customer behavior variables as nodes, filter the causal edges according to the structural correlation, and form a directed behavior path graph; Use behavior disturbance test method to change the behavior frequency of variables in the path, and calculate the change of loss probability distribution before and after the change; Take the change of loss probability caused by behavior disturbance as the path effectiveness criterion, and select the high sensitivity path as the core early warning path set. 4.The insurance customer churn early warning method of fusing causal inference learning mechanism according to claim 1, characterized in that, The determination of customer loss risk in S3 includes: Based on the current node state of the customer in the high causal degree path, combine the historical evolution trend of the path, and predict the possibility of developing to non-renewal state; If the similarity between the current behavior of the customer and the typical loss customer behavior in the early warning path is higher than the set threshold, it is determined that there is a potential loss risk. 5.The insurance customer churn early warning method of fusing causal inference learning mechanism according to claim 1, characterized in that, The risk level determination rule in S4 includes: If the predicted churn probability is greater than a preset second threshold, the customer is marked as a high-risk customer; If the prediction result is between the first threshold and the second threshold, the customer is marked as a medium-risk customer; Otherwise, the customer is marked as a non-risk customer; The thresholds can be automatically calibrated based on historical data or set by operational management rules.

6. The insurance customer churn early warning method fusing causal inference learning mechanism according to claim 1, characterized in that, The calibration of the behavior intervention point in S4 includes: According to the comparison of the customer's current behavior node and the early warning path structure, identify the path node with the possibility of behavior reversal and mark it as an intervenable node; If a behavior sequence is interrupted and the customer has no active signal or the access path is completely consistent with the risk customer, it is marked as an irreversible node; Output the two types of nodes respectively for personalized intervention recommendation and risk closed-loop management.

7. The insurance customer churn early warning method fusing causal inference learning mechanism according to claim 1, characterized in that, The artificial intervention suggestion pushing in S5 includes: Generate a risk profile for each customer, including customer identification, behavior label, risk level, and core path node information; Generate corresponding intervention suggestion scripts for each type of customer based on their historical renewal status and current behavior status; Synchronize the high-risk customer list to the sales, customer service, and agent channel systems to trigger exclusive communication or task reminder processes.

8. An insurance customer churn early warning system fusing causal inference learning mechanism, for implementing the method of any one of claims 1-7, characterized in that, It includes: A customer data collection module for synchronizing multi-source customer behavior data from customer management systems, policy management systems, and online platforms; A causal path analysis module for establishing a path structure diagram based on behavior variables and calculating causal scores; A risk state reasoning module for predicting future renewal states and classifying risk levels; A node identification and intervention suggestion module for generating key behavior nodes and corresponding early warning suggestion information; A risk output and task linkage module for exporting high-risk customer data to external intervention systems or artificial intervention platforms.

Citation Information

Patent Citations

  • Medical adverse event risk prediction, prevention and control method based on causal inference

    CN119993487A

  • Insurance product intelligent pushing mode based on deep neural network

    CN120013629A

  • Data operation system and method based on knowledge graph

    CN120705496A

  • Integrated computer software and hardware combined sales optimization system and method

    CN120765297A

  • Multi-modal causal reasoning and explaining method, device, equipment and medium

    CN120952184A