Group insurance underwriting method based on artificial intelligence
By employing an AI-based group insurance underwriting method that utilizes multi-source data preprocessing and a multi-model fusion engine, the problem of low efficiency in traditional group insurance underwriting is solved. This enables efficient and accurate automated underwriting decisions, reducing operating costs and the risk of data leakage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional group insurance underwriting faces the challenge of processing massive amounts of heterogeneous data, resulting in long underwriting times and low accuracy. Reliance on manual labor leads to inefficiency and high costs.
An AI-based group insurance underwriting method is adopted, which uses multi-source data preprocessing, composite risk feature set construction, and Stacking integrated learning architecture to build a multi-model fusion engine for risk assessment and prediction, and combines a rule engine to achieve automated decision-making.
Significantly shorten underwriting time, improve accuracy and consistency, reduce operating costs, enhance customer experience and profitability, reduce the risk of data breaches, and support model optimization and regulatory compliance.
Smart Images

Figure CN121746090A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and insurance underwriting technology, specifically to an artificial intelligence-based group insurance underwriting method. Background Technology
[0002] With the continuous improvement of my country's social security system and the rapid development of the commercial insurance market, group insurance (hereinafter referred to as "group policy"), as a core component of corporate employee benefits, has seen its market size continue to expand. According to the 2022 annual report of the Insurance Association of China, the annual premium scale of group policy business has exceeded 100 billion yuan, with an average annual growth rate of over 15%. Group policy business is characterized by high average premiums per policy and strong customer loyalty, and has become a key business area for major insurance companies. However, compared with traditional individual insurance underwriting, group policy underwriting faces more complex risk assessment and management challenges. Currently, the group policy underwriting technology commonly used in the industry mainly adopts a manual underwriting model. Underwriters need to process massive amounts of heterogeneous insurance data, including but not limited to: basic information of the insured company (industry, size, region, operating status), claims data over the years, employee demographic characteristics (age, gender distribution), occupational categories, health disclosure information, etc. According to the "Insurance Underwriting Efficiency Research Report" (2023), a senior underwriter takes an average of 4-6 working hours to process a group insurance application for a medium-sized enterprise (200-500 employees), and even longer for complex cases. Against this backdrop, it is necessary to consider ways to improve underwriting efficiency and accuracy. Summary of the Invention
[0003] The purpose of this application is to provide a group insurance underwriting method based on artificial intelligence, and the specific technical solution is as follows:
[0004] An AI-based group insurance underwriting method includes: S1, acquiring and preprocessing multi-source data information from the insured company at both the enterprise and employee levels; S2, constructing a composite risk feature set at both the enterprise and employee levels based on the preprocessed data in S1; S3, evaluating the correlation between the preprocessed data in S1 and the composite risk feature set constructed in S2, selecting highly correlated features above a preset threshold, and generating feature vectors accordingly; S4, conducting a preliminary review of the feature vectors generated in S3 to screen for policies that are clearly compliant or non-compliant; S5, constructing a multi-model fusion engine for risk prediction using a Stacking ensemble learning architecture; and S6, using the multi-model fusion engine constructed in S5 to perform risk scoring on the compliant policies screened in S4 and outputting the final underwriting conclusion.
[0005] Preprocessing in S1 includes: S1.1, receiving multi-source data streams in real time, including structured and unstructured data uploaded by enterprises, as well as credit and medical data; S1.2, performing integrity verification, missing value processing, and outlier detection on the data received in S1.1, and uniformly converting heterogeneous data into a standard format.
[0006] When constructing the composite risk feature set at the enterprise and employee levels in S2, the following steps are included: S2.1: Analyzing the characteristics of enterprise operations, industry risks, and historical underwriting and claims information to construct enterprise-level features, establishing an enterprise operational stability assessment system and an industry risk coefficient calculation model, integrating indicators such as the enterprise's establishment year, revenue growth stability, employee turnover rate, and debt-to-equity ratio, and comprehensively considering the industry's historical insurance products, loss ratio, and regional differences. Each indicator is weighted and integrated after standardization. S2.2: For individual employee risks, a health risk scoring model is established. This model is based on historical claims data to statistically calculate the weight coefficients of various diseases, combined with abnormal physical examination indicators and family medical history information for comprehensive scoring, constructing an occupational risk level classification system, and forming an occupational risk matrix.
[0007] When generating feature vectors in S3, the following steps are taken: the preprocessed data in S1 is used as the basis for the composite risk feature set constructed in S2, and the feature importance is ranked based on XGBoost. Features with importance scores > 0.01 are retained, and finally, feature vectors are generated.
[0008] The multi-model fusion engine for risk prediction in S5 is implemented using a two-layer model structure, specifically including: S5.1, building base learners based on XGBoost, LightGBM, random forest, and neural network models. These base learners learn from the original training data and generate prediction results. The diversity of base learners ensures that the model can capture risk patterns in the data from different perspectives; S5.2, using the prediction results of the base learners in S5.1 as input features, and using a logistic regression model as a meta-learner, whose input is the prediction results of the base learners, and whose output is the final risk score.
[0009] The multi-model fusion engine in S5 employs K-fold cross-validation during training. K-fold cross-validation (K=5) is used to generate predictions for the base learners. Specifically, this includes: S5.3, dividing the training set into 5 folds, performing 5 training and prediction iterations for each base learner, using 4 folds for training and predicting the remaining 1 fold each time to obtain the prediction results for the entire training set; simultaneously, performing 5 predictions on the test set and averaging the results; S5.4, using the prediction results obtained in S5.3 as training data for the meta-learner, and employing L2 regularization in the meta-learner to further control overfitting to prevent it.
[0010] The performance optimization and validation of the multi-model fusion engine in S5 includes: S5.5, using Bayesian optimization algorithm to tune the hyperparameters of the base learner and meta-learner, with the search space including learning rate, tree model depth, and regularization strength parameters; S5.6, the optimization objective is the AUC value on the validation set, and the optimal parameter combination is found through multiple rounds of iteration; S5.7, a three-level model validation system is established, with the technical aspect evaluating the model's discrimination ability through AUC, KS value, and F1 score indicators, the business aspect verifying business applicability through payout deviation rate and risk discrimination indicators, and the compliance aspect conducting fairness verification and interpretability audit.
[0011] S6 includes: a comprehensive decision based on the prediction results of the multi-model fusion engine built in S5 and business rules, which is transformed into underwriting decisions through risk quantification units to provide a basis for decision-making, thereby generating the final underwriting conclusion and pushing the underwriting results to the insured enterprise and the business side of the insurance company, and synchronizing with the core business system to generate insurance policies.
[0012] The beneficial effects of this application are as follows: AI-generated underwriting results significantly shorten underwriting time and improve efficiency. The automated processing mechanism reduces human intervention, minimizes human error, and enhances the accuracy and consistency of underwriting. The underwriting model combines customer claims history, product past claims models, and professional risk models, enabling a comprehensive and accurate risk assessment. Through big data analysis and AI technology, the system can more accurately identify high-risk customers, reducing the operational risk for insurance companies. The automated processing mechanism reduces the workload of manual review, lessens the absolute reliance on senior underwriters, and lowers operating costs. Simultaneously, the efficient underwriting process improves customer experience, and more accurate risk identification and fraud prevention directly reduce insurance company payouts, improving overall profitability and risk control. It establishes a new underwriting work model, freeing human experts from tedious and repetitive tasks, allowing them to focus on handling the most complex and challenging edge cases. It supports underwriting experts' feedback and optimization of the AI model, improving the overall quality and efficiency of underwriting operations. Through a closed-loop feedback learning mechanism, the system's AI model continuously learns from underwriting experts' decisions and actual claims results, adapting to changes in market risks and continuously optimizing underwriting strategies. Sensitive customer data is encrypted and stored, effectively protecting customer information. Underwriting conclusions are exported directly to a cloud drive, allowing operators to process and manage data on the cloud drive, with automatic data file cleanup fundamentally eliminating the risk of data leakage. Complete operation audit logs meet regulatory compliance requirements. Multi-dimensional analysis of underwriting data and regular generation of risk assessment reports not only classify the risk of insured companies and customers but also provide in-depth data analysis, enabling dynamic risk warning mechanisms and data-driven product pricing strategies. This provides data support for decision-making and management, improving the company's operational capabilities and profitability. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating the application process. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to specific embodiments and accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of this application. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0015] like Figure 1 As shown, an artificial intelligence-based group insurance underwriting method includes:
[0016] S1. Obtain and preprocess multi-source data information submitted by the insured company at both the enterprise and employee levels. Specifically, the preprocessing includes: S1.1. Receiving multi-source data streams in real time, including structured data uploaded by the company (financial information of the insured company, employee list) and unstructured data (business license, employee experience report), as well as accessing credit and medical data; S1.2. Performing integrity verification, missing value handling, and outlier detection on the data received in S1.1, and uniformly converting heterogeneous data into a standard format.
[0017] S2. Based on the preprocessed data in S1, construct a composite risk feature set at the enterprise and employee levels. Specifically, constructing the composite risk feature set at the enterprise level (industry risk coefficient, employee size risk) and the employee level (health risk score, occupational risk level) includes: S2.1. Analyzing the characteristics of enterprise operations, industry risks, and historical underwriting and claims information to construct enterprise-level features, establishing an enterprise operational stability assessment system and an industry risk coefficient calculation model, integrating indicators such as the enterprise's establishment year, revenue growth stability, employee turnover rate, and debt-to-equity ratio, comprehensively considering the industry's historical insurance products, loss ratio, and regional differences, and weighting and integrating each indicator after standardization; S2.2. For individual employee risks, establish a health risk scoring model. This model is based on historical claims data to statistically calculate the weight coefficients of various diseases, combined with abnormal physical examination indicators and family medical history information for comprehensive scoring, constructing an occupational risk level classification system, and forming an occupational risk matrix.
[0018] S3. The preprocessed data from S1 is used to evaluate the relevance of the composite risk feature set constructed in S2, and highly relevant features with a value higher than a preset threshold are selected to generate a feature vector. Specifically, generating the feature vector includes: using the preprocessed data from S1 and the composite risk feature set constructed in S2, the feature importance is ranked based on XGBoost, and features with an importance score > 0.01 are retained, ultimately generating the feature vector.
[0019] S4. Conduct a preliminary review of the feature vectors generated in S3 to screen for policies that are clearly compliant or non-compliant.
[0020] S5. A multi-model fusion engine for risk prediction is constructed using a Stacking ensemble learning architecture. Specifically, the multi-model fusion engine for risk prediction is implemented using a two-layer model structure, including: S5.1, building base learners based on XGBoost, LightGBM, random forest, and neural network models. These base learners learn from the original training data and generate prediction results. The diversity of the base learners ensures that the model can capture risk patterns in the data from different perspectives; S5.2, using the prediction results of the base learners in S5.1 as input features, a logistic regression model is used as the meta-learner, with the base learner's prediction results as input and the final risk score as output.
[0021] The multi-model fusion engine employs K-fold cross-validation during training. K-fold cross-validation (K=5) is used to generate predictions for the base learners. Specifically, this includes: S5.3, dividing the training set into 5 folds, and performing 5 training and prediction iterations for each base learner. Each iteration uses 4 folds for training and predicts the remaining 1 fold, thus obtaining the prediction results for the entire training set. Simultaneously, 5 predictions are performed on the test set, and the average value is taken. S5.4, using the prediction results obtained in S5.3 as the training data for the meta-learner, the meta-learner is trained. To prevent overfitting, L2 regularization is applied in the meta-learner to further control overfitting.
[0022] The performance optimization and validation of the multi-model fusion engine includes: S5.5, using Bayesian optimization algorithm to tune the hyperparameters of the base learner and meta-learner, with the search space including learning rate, tree model depth, and regularization strength parameters; S5.6, the optimization objective is the AUC value on the validation set, and the optimal parameter combination is found through multiple iterations; S5.7, a three-level model validation system is established: at the technical level, the model's discrimination ability is evaluated through AUC, KS value, and F1 score; at the business level, the business applicability is verified through payout deviation rate and risk discrimination index; and at the compliance level, fairness verification and interpretability audit are conducted.
[0023] S6. The multi-model fusion engine built in S5 is used to perform risk scoring on the compliant policies screened in S4, and output the final underwriting conclusion. Specifically, this includes: a comprehensive decision based on the prediction results of the multi-model fusion engine built in S5 and business rules, which is transformed into an underwriting decision through a risk quantification unit to provide a basis for decision-making, thereby generating the final underwriting conclusion and pushing the underwriting results to the insured company and the business side of the insurance company, and synchronizing with the core business system to generate policies.
[0024] In practical applications, static underwriting rules (such as industry premium standards, age limits, and maximum coverage limits) can be entered through a visual interface. This supports personalized rule configuration and adjustment, condition combinations, and priority settings for group insurance business, and provides historical version management functions for rules. At the same time, by setting up a collaborative strategy between the AI model and the rule engine, a division of labor and cooperation mechanism with the AI model can be realized (such as triggering the AI intelligent underwriting engine after the initial review of the rules is passed).
[0025] The group booking AI underwriting device proposed in this application includes the following core modules, which interact in real time via a data bus:
[0026] Data access and preprocessing module: Supports multi-source access of structured data (company information, employee list, claims records) and unstructured data (medical examination reports, medical texts), with built-in data cleaning engine (duplicate removal, completion, format standardization) and OCR+NLP text parsing unit to extract risk features (such as medical history, abnormal medical examination indicators) from unstructured data.
[0027] Underwriting rule engine module: includes a static rule base (industry rate standards, disclaimers) and a dynamic rule generation unit, supports visual rule configuration, and can work with AI models to trigger underwriting decisions (such as rule preliminary review + AI intelligent underwriting).
[0028] The AI-powered intelligent underwriting module consists of a feature engineering unit, a multi-model fusion engine, and a risk quantification unit. The feature engineering unit constructs multi-dimensional risk indicators (at the enterprise level: industry risk, operational stability; at the employee level: age distribution, health risk, occupational risk). The multi-model fusion engine integrates XGBoost and neural network models through a stacking architecture, achieving an AUC value ≥ 0.85, thus enabling accurate prediction of risk probabilities. The risk quantification unit transforms the prediction results into actionable conclusions such as underwriting rates and exclusions.
[0029] Decision Collaboration and Output Module: Establishes a two-layer decision-making mechanism of "rule engine + AI model" to automatically approve low-risk policies, reject high-risk policies, and transfer medium-risk policies to manual review; supports the visual export of underwriting conclusions (including risk point explanations and premium calculation basis), and connects with core business systems to achieve data synchronization.
[0030] Closed-loop feedback module: Collects subsequent claims data and manual underwriting correction records of the policy, and automatically feeds them back to the AI model training unit to achieve weekly iterative optimization of the model and adapt to changes in risk characteristics.
[0031] To make this application easier to understand, the following explanation is further illustrated with practical application examples.
[0032] Hardware environment configuration
[0033] Server configuration: CPU is Intel Xeon Gold 6330 (≥2), memory is ≥128GB, hard drive is ≥2TB SSD, GPU is NVIDIA A100 (≥1 unit, used for AI model training and inference).
[0034] Network environment: Supports Gigabit Ethernet, latency ≤5ms, ensuring real-time access to multi-source data and system integration.
[0035] Storage environment: A distributed database (Hadoop + MySQL) is used. Structured data is stored in MySQL, while unstructured data and feature data are stored in Hadoop, supporting data backup and fast retrieval.
[0036] Software environment configuration
[0037] Operating system: CentOS 7.9 64-bit.
[0038] Development languages: Python 3.9 (AI model development), Java 11 (business system development).
[0039] Core framework and tools: HanLP is used for NLP parsing, TensorFlow 2.8+Scikit-learn is used for machine learning framework, Drools is used for rule engine, Spark is used for data processing, and ECharts is used for visualization tool.
[0040] Specific implementation steps
[0041] The main workflow is as follows:
[0042] Configure underwriting rules
[0043] Underwriters can input static rules through a visual interface, such as "10% surcharge for manufacturing employees whose average age is >45" and "exclusion of related liabilities for those diagnosed with malignant tumors," and set the collaborative strategy to "trigger AI intelligent underwriting after the rule is approved in the initial review."
[0044] Insurance data access and preprocessing
[0045] Users enter group insurance applications through the front-end interface, including the insurance plan number, business license (structured), employee list (Excel), and medical examination report (PDF). The data access module automatically verifies data integrity, the cleaning engine removes duplicate employee records, and the NLP parsing unit extracts risk features such as "Hypertension Grade III" and "Abnormal Blood Sugar" from the medical examination report, converting them into structured indicators.
[0046] Feature engineering processing
[0047] Based on indicators such as enterprise industry (e.g., the risk coefficient for the construction industry is set to 1.2), employee age distribution (calculating average age and the proportion of employees over 45 years old), and health risk score (assigned according to the severity of disease), a composite feature set is formed for model input, and the top 10 highly relevant features such as "industry risk coefficient, average age, and health risk score" are automatically selected.
[0048] AI-powered intelligent underwriting decision
[0049] Preliminary review by the rules engine: A group purchase from a manufacturing company (300 employees, average age 42) has no violations, good corporate credit, and a controllable historical claims ratio, and passes the preliminary review. AI model underwriting: The multi-model fusion engine calls a pre-trained model, inputs a composite feature set, outputs a risk level of "low risk," and the underwriting conclusion is "standard rate approved." Abnormal case handling: In a group purchase from a technology company, 5 employees were diagnosed with diabetes. The rules engine triggers the "specific disease exclusion" rule. The AI model calculates additional risk coefficients and concludes "excluding diabetes-related liabilities, standard rate approved." Manual review: Underwriting experts receive manual underwriting tasks on the review task interface, review all materials provided by the system (raw data, risk scores, risk tags), make final underwriting review opinions, and record them in the system, building an expert review collaborative workflow to achieve intelligent decision-making through human-machine collaboration.
[0050] Underwriting conclusion output and system integration
[0051] Establish a data security protection mechanism, generate underwriting conclusions based on intelligent underwriting, and support exporting to cloud storage or pushing to the insured company's email address. All exported data is anonymized to ensure that the data is not leaked, and all processing traces are recorded for later auditing and traceability. At the same time, it is synchronized to the insurance company's core business system to generate a policy number.
[0052] Closed-loop feedback and model optimization
[0053] The system collects claims data for underwritten policies monthly. If the actual claims rate for a certain industry is 15% higher than the predicted value, the closed-loop feedback module automatically incorporates the data for that industry into the model training, adjusts the feature weights, optimizes the risk prediction algorithm, and prompts underwriting experts to update relevant rules.
[0054] Implementation effect verification
[0055] Efficiency Improvement: Underwriting time for group policies of medium-sized enterprises has been reduced from 4-6 hours to ≤30 minutes, the average daily processing volume has increased from 2-3 policies to ≥20 policies, and the policy issuance cycle has been shortened from 5-10 working days to 1-2 working days; Accuracy Improvement: The AI model AUC value is stable at 0.88-0.92, the claims rate prediction deviation is ≤8%, and the underwriting rate difference rate is ≤3% (unified standard); Cost Reduction: The proportion of manual underwriting has decreased from 70% to 20%, and underwriting operating costs have been reduced by more than 40%; Customer Experience: The self-underwriting approval rate has increased from 75% to 85%, the insurance progress can be tracked in real time, and the number of plan revisions has been reduced by ≥60%.
Claims
1. A group insurance underwriting method based on artificial intelligence, characterized in that, include: S1. Obtain and preprocess multi-source data information from the enterprise and employee levels submitted by the insured company. S2. Construct a composite risk feature set at the enterprise level and the employee level based on the preprocessed data in S1; S3. The preprocessed data in S1 is evaluated for correlation based on the composite risk feature set constructed in S2, and highly correlated features above a preset threshold are selected to generate feature vectors. S4. Conduct a preliminary review of the feature vectors generated in S3 to screen for insurance policies that are obviously compliant or non-compliant; S5. A multi-model fusion engine for risk prediction is built using a Stacking ensemble learning architecture. S6. Use the multi-model fusion engine built in S5 to perform risk scoring on the compliant insurance policies screened in S4, and output the final underwriting conclusion.
2. The group insurance underwriting method based on artificial intelligence as described in claim 1, characterized in that, The preprocessing in S1 includes: S1.1 Real-time reception of multi-source data streams, including structured and unstructured data uploaded by enterprises, as well as access to credit and medical data; S1.2 Perform integrity verification, missing value processing, and outlier detection on the data received in S1.1, and convert heterogeneous data into a standard format.
3. The group insurance underwriting method based on artificial intelligence as described in claim 2, characterized in that, The construction of the composite risk feature set at the enterprise and employee levels in S2 includes: S2.1 Analyze the characteristics of enterprise operation, industry risk, historical underwriting and claims information, construct enterprise-level characteristics, establish an enterprise operation stability assessment system and industry risk coefficient calculation model, integrate indicators of enterprise establishment years, revenue growth stability, employee turnover rate and asset-liability ratio, comprehensively consider the industry's historical insurance products, loss ratio and regional difference factors, and weight and merge the indicators after standardization. S2.
2. For individual employee risks, establish a health risk scoring model. This model is based on the weight coefficients of various diseases calculated from historical claims data, combined with abnormal physical examination indicators and family medical history information to conduct a comprehensive score, construct an occupational risk level classification system, and form an occupational risk matrix.
4. The group insurance underwriting method based on artificial intelligence as described in claim 3, characterized in that, The process of generating feature vectors in S3 includes: taking the preprocessed data in S1 and the composite risk feature set constructed in S2, sorting the features by importance based on XGBoost, retaining features with importance scores > 0.01, and finally generating feature vectors.
5. The group insurance underwriting method based on artificial intelligence as described in claim 4, characterized in that, The S5 uses a two-layer model structure to build a multi-model fusion engine for risk prediction, specifically including: S5.
1. Base learners are built based on XGBoost, LightGBM, random forest and neural network models. These base learners learn from the original training data and produce prediction results. The diversity of base learners ensures that the model can capture risk patterns in the data from different perspectives. S5.
2. Using the prediction results of the base learner in S5.1 as input features, a logistic regression model is used as the meta-learner, with the prediction results of the base learner as input and the final risk score as output.
6. The group insurance underwriting method based on artificial intelligence as described in claim 5, characterized in that, The multi-model fusion engine in S5 employs K-fold cross-validation during training, using K-fold cross-validation (K=5) to generate predictions for the base learner, specifically including: S5.3 Divide the training set into 5 folds. For each base learner, perform 5 training and predictions. Each time, use 4 folds for training and predict the remaining 1 fold to obtain the prediction results for the entire training set. At the same time, perform 5 predictions on the test set and take the average value. S5.
4. Use the prediction results obtained in S5.3 as the training data for the meta-learner to train the meta-learner. To prevent overfitting, L2 regularization is used in the meta-learner to further control overfitting.
7. The group insurance underwriting method based on artificial intelligence as described in claim 6, characterized in that, The performance optimization and verification of the multi-model fusion engine in S5 includes: S5.
5. The hyperparameters of the base learner and meta learner are tuned using the Bayesian optimization algorithm. The search space includes the learning rate, tree model depth, and regularization strength parameter. S5.6 The optimization objective is the AUC value on the validation set, and the optimal parameter combination is found through multiple rounds of iteration; S5.7 Establish a three-level model validation system. At the technical level, the model's distinguishing ability is evaluated through AUC, KS value and F1 score indicators. At the business level, the business applicability is verified through the payout deviation rate and risk distinguishability indicators. At the compliance level, fairness is tested and interpretability is audited.
8. The group insurance underwriting method based on artificial intelligence as described in claim 7, characterized in that, S6 includes: making a comprehensive decision based on the prediction results of the multi-model fusion engine constructed in S5 and the business rules, transforming it into an underwriting decision through a risk quantification unit, providing a basis for decision-making, thereby generating a final underwriting conclusion and pushing the underwriting results to the insured enterprise and the business end of the insurance company, and synchronizing it to the core business system to generate an insurance policy.