Insurance customer cross-institution credit risk assessment system based on federal learning
By constructing a federated learning system that extracts risk features specific to the insurance industry, protects privacy, and incentivizes cross-institutional collaboration, the problems of insufficient targeting of assessment models and inaccurate privacy protection in existing technologies have been solved, achieving an efficient, secure, and feasible integrated solution for credit risk assessment.
Patent Information
- Application Number
- CN202511676967.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing federated learning-based insurance credit risk assessment solutions fail to build an integrated system that combines insurance industry-specific risk feature extraction, data sensitivity classification and privacy protection, and cross-institutional collaboration incentives for insurance scenarios. This results in assessment models that are not targeted enough, lack accuracy in privacy protection, and have weak motivation for cross-institutional collaboration.
It provides a cross-institutional credit risk assessment system for insurance customers based on federated learning, including a data preprocessing module, a federated training module, a privacy protection module, a blockchain incentive module, an interpretability enhancement module, and a model optimization module. Through secure alignment, feature extraction, privacy protection, cross-institutional incentives, and model optimization, it improves the relevance and privacy security of the assessment system.
It significantly improves the comprehensiveness and accuracy of credit risk assessment, alleviates model bias caused by uneven data distribution, realizes an incentive mechanism for cross-institutional collaboration, and meets the business scenario adaptability and regulatory compliance needs of the insurance industry.
Smart Images

Figure CN121504632A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of credit risk assessment, and particularly relates to an insurance customer cross-institution credit risk assessment system based on federated learning. BACKGROUND
[0002] In the process of insurance business, credit risk assessment is the core link of underwriting pricing, anti-fraud detection and reinsurance arrangement, and its accuracy directly affects the stability of industry operation. With the diversification development of insurance business, the limited data dimension within a single institution, cross-institution data fusion becomes the key to improve the evaluation accuracy, but the strengthening of data privacy protection regulations restricts the direct sharing of original data, forming the "data island" problem. As a distributed learning paradigm of "data available but invisible", federated learning provides technical support for cross-institution cooperation, and promotes the innovation of insurance credit risk assessment technology.
[0003] The existing evaluation scheme based on federated learning fails to construct an integrated system integrating insurance industry-specific risk feature extraction, data sensitivity grading privacy protection and cross-institution cooperation incentive for the insurance scene, resulting in insufficient pertinence of the evaluation model, lack of privacy protection accuracy and weak cross-institution cooperation motivation, which restricts the improvement of the efficiency of cross-institution credit risk assessment of insurance customers. Therefore, it is necessary to provide an insurance customer cross-institution credit risk assessment system based on federated learning to solve the above technical problems. SUMMARY
[0004] The present application provides an insurance customer cross-institution credit risk assessment system based on federated learning, which solves the problem that the existing evaluation scheme based on federated learning fails to construct an integrated system integrating insurance industry-specific risk feature extraction, data sensitivity grading privacy protection and cross-institution cooperation incentive for the insurance scene, resulting in insufficient pertinence of the evaluation model, lack of privacy protection accuracy and weak cross-institution cooperation motivation.
[0005] To solve the above technical problems, the insurance customer cross-institution credit risk assessment system based on federated learning provided by the present application comprises: A data preprocessing module is used for secure alignment and standardized processing of cross-institution multi-source data, generating a feature data set available for federated training, and an insurance industry-specific risk feature extraction unit is built in, which is used for automatically generating insurance industry-specific risk features for policy data, claim records and credit information; A federated training module is used for distributed model training based on multi-party feature data sets, generating a global credit risk assessment model through a dynamic aggregation strategy, and supporting a hierarchical training mechanism based on insurance risk levels, and using a reinforcement learning strategy to optimize the model convergence direction for high-risk customer samples; The privacy protection module is used to protect the privacy of data, model parameters and intermediate results during training, and has a built-in risk data sensitivity classification unit to dynamically adjust the privacy protection strength according to the sensitivity level of insurance data. The blockchain incentive module is used to quantify the contribution of cross-institutional data and computing resources, realize automated incentive allocation, and its contribution measurement rules are bound to insurance business scenarios. The interpretability enhancement module is used to generate the logical basis and visual explanation results for model decisions, and supports the output of explanations for insurance regulatory compliance. The model optimization module is used to dynamically adjust the model structure and training strategy to improve the model's generalization ability, and it also has a built-in insurance business cycle adaptation unit to dynamically optimize the model update frequency according to business characteristics. The risk assessment module is used to output the credit risk level and assessment results of insurance customers based on a global model, and supports the linked output of risk level and insurance product pricing rules.
[0006] As a preferred embodiment of this application, the data preprocessing module includes a security ID alignment unit, a multimodal feature unified representation unit, a feature standardization unit, and an insurance industry-specific risk feature extraction unit; The secure ID alignment unit uses privacy set intersection technology to achieve de-identification matching of cross-institutional customer identities and supports multi-dimensional cross-verification of insurance customer unique identifiers. The multimodal feature unified representation unit transforms textual, numerical, and time-series insurance credit-related data into feature vectors of a unified dimension through a multimodal large model; The feature standardization unit uses a local independent processing method to complete feature cleaning and normalization transformation; The insurance industry-specific risk feature extraction unit automatically generates insurance industry-specific risk features based on actuarial rules. As a preferred embodiment of this application, the federated training module supports adaptive switching between vertical federated learning and horizontal federated learning, and includes hierarchical training units based on insurance risk levels. The vertical federated learning is suitable for scenarios with overlapping samples and complementary features across institutions, and generates cross-institutional cross features through secure multi-party computation; The horizontal federated learning is suitable for scenarios where cross-institutional features are consistent and samples are complementary, and an improved federated averaging algorithm is used to aggregate model parameters. The stratified training unit based on insurance risk level stratifies customer samples according to preset risk levels and uses enhanced training with a higher number of iterations for high-risk samples. The federated training module also includes a large model-assisted pre-training unit and a global alignment unit. The large model-assisted pre-training unit uses a server-side multimodal large language model for global pre-training, and the global alignment unit performs global semantic alignment on the locally fine-tuned model based on the pre-trained model.
[0007] As a preferred embodiment of this application, the federated training module adopts a dynamic adaptive aggregation algorithm, which dynamically adjusts the aggregation weights of the participating parameters based on the data distribution characteristics, data quality and model performance of each participating institution, and assigns weighting coefficients to institutions containing high-value risk characteristics, so as to alleviate the model bias problem caused by non-independent and identically distributed data.
[0008] As a preferred embodiment of this application, the privacy protection module adopts a hybrid privacy protection mechanism that combines homomorphic encryption, differential privacy and trusted execution environment, and includes a risk data sensitivity classification unit; The risk data sensitivity classification unit divides insurance data into core sensitive data, general sensitive data, and non-sensitive data, and configures different encryption strengths accordingly. The homomorphic encryption is used to support feature calculation and parameter aggregation in the dense state; the differential privacy is used to add noise perturbation to the uploaded model gradient or parameters, and the noise intensity is dynamically adjusted according to the data sensitivity. The trusted execution environment is used to construct a hardware-level secure computing space, isolating sensitive data and computational logic during the training process.
[0009] As a preferred embodiment of this application, the blockchain incentive module is implemented based on a consortium blockchain architecture, including a contribution measurement unit and a smart contract incentive unit; The contribution measurement unit comprehensively scores the data provided by participating institutions based on the amount of data, data quality, computing resource input, and model performance contribution, and assigns extra weight to data that provides key anti-fraud features. The smart contract incentive unit automatically executes the incentive allocation rules based on the contribution score, generates an immutable incentive record and stores it on the blockchain, and supports integration with the incentive settlement interface of the insurance business system.
[0010] As a preferred embodiment of this application, the interpretability enhancement module includes a local decision interpretation unit, a global aggregation visualization unit, and a regulatory compliance interpretation unit; The local decision interpretation unit uses a decision tree model to match the correlation logic between credit risk assessment results and input features, and outputs the key influencing factors of a single customer's risk score; The global aggregation visualization unit generates a visual view of the model aggregation process through feature importance analysis and dimensionality reduction techniques. The regulatory compliance interpretation unit automatically generates a structured report containing model principles, feature selection criteria, and risk level classification standards in accordance with insurance regulatory requirements, thus meeting the disclosure requirements of insurance regulatory agencies.
[0011] As a preferred embodiment of this application, the model optimization module includes an incremental learning unit, a model compression unit, an edge computing adaptation unit, and an insurance business cycle adaptation unit. The incremental learning unit supports dynamic model updates based on new data, without the need to retrain the global model. The model compression unit reduces model storage and transmission overhead through knowledge distillation and parameter quantization techniques; The edge computing adaptation unit is used to adapt the global model to the edge node to achieve rapid local risk assessment; The insurance business cycle adaptation unit dynamically adjusts the model update frequency based on the characteristics of the insurance product cycle and claims processing time.
[0012] This invention also provides a method for cross-institutional credit risk assessment of insurance clients based on the above system, the method comprising the following steps: Step 1: Cross-institutional participants complete the local data security ID alignment, multimodal feature unified representation, standardization processing, and insurance industry-specific risk feature extraction through the data preprocessing module to generate a local feature dataset; Step 2: The federated training module receives training requests from each participant, adaptively selects vertical or horizontal federated learning mode based on data distribution characteristics, and conducts local model training and parameter updates in layers according to insurance risk level based on the risk data sensitivity classification protection mechanism of the privacy protection module. Before training, global pre-training is completed through the server-side multimodal large language model, and then the pre-trained model is distributed to each participant for local fine-tuning. After fine-tuning, semantic alignment is achieved through the global alignment unit. Step 3: The federated training module uses a dynamic adaptive aggregation algorithm to globally aggregate the encrypted model parameters uploaded by each participant, assign weighting coefficients to high-value risk characteristic institutions, and generate an initial global model. Step 4: The interpretability enhancement module performs decision logic analysis, visualization processing, and regulatory compliance report generation on the initial global model, outputting a model interpretability report; Step 5: The model optimization module, in conjunction with the characteristics of the insurance business cycle, performs incremental updates, compression, and edge adaptation optimization on the initial global model to generate the final global credit risk assessment model. Step Six: Based on the final global model, the risk assessment module classifies the credit risk of the target insurance customers and outputs assessment results that match the pricing rules of the insurance products. Step 7: The blockchain incentive module automatically completes the incentive distribution and on-chain storage based on the contributions of each participant during the training process through smart contracts.
[0013] Compared with related technologies, the cross-institutional credit risk assessment system for insurance customers based on federated learning provided by this invention has the following beneficial effects: 1. This invention relies on a distributed training mode where federated learning data is available but not visible, coupled with a hierarchical encryption mechanism for privacy protection modules. While meeting data privacy compliance requirements, it quantifies institutional contributions and automatically allocates incentives through a blockchain incentive module, fully mobilizing the enthusiasm for cross-institutional collaboration, and integrating multi-dimensional data such as insurance policies, claims, and credit reports to significantly improve the comprehensiveness of credit risk assessment.
[0014] 2. This invention extracts insurance-specific risk features through a data preprocessing module and combines it with a high-risk stratified training mechanism of federated training to accurately capture industry-specific risk patterns. The incremental update and business cycle adaptation functions of the model optimization module allow the model to dynamically adapt to business characteristics such as the insurance period and claims processing time without full retraining, effectively mitigating model bias caused by uneven data distribution.
[0015] In summary, this invention specifically addresses the core issues in cross-institutional credit risk assessment in the insurance industry, such as difficulties in data collaboration, poor model adaptability, inaccurate privacy protection, insufficient compliance, and weak motivation for collaboration. Through multi-module collaboration, it achieves a comprehensive improvement in the accuracy of risk assessment, adaptability to business scenarios, privacy and security protection, and regulatory compliance, providing an efficient, secure, and feasible integrated technical solution for cross-institutional credit risk assessment in the insurance industry. Attached Figure Description
[0016] Figure 1 This is a block diagram illustrating the connection principle of each module in the cross-institutional credit risk assessment system for insurance customers based on federated learning provided by this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “group,” “class,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0019] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0020] Please refer to the following: Figure 1 A federated learning-based cross-institutional credit risk assessment system for insurance clients includes: The data preprocessing module is used for secure alignment and standardization of multi-source data across institutions, generating feature datasets usable for federated training. It also has a built-in risk feature extraction unit specifically for the insurance industry, which is used to automatically generate insurance industry-specific risk features for policy data, claims records, and credit information. The federated training module is used to conduct distributed model training based on multi-party feature datasets. It generates a global credit risk assessment model through dynamic aggregation strategies and supports a hierarchical training mechanism based on insurance risk levels. For high-risk customer samples, a reinforcement learning strategy is used to optimize the model convergence direction. The privacy protection module is used to protect the privacy of data, model parameters and intermediate results during training, and has a built-in risk data sensitivity classification unit to dynamically adjust the privacy protection strength according to the sensitivity level of insurance data. The blockchain incentive module is used to quantify the contribution of cross-institutional data and computing resources, realize automated incentive allocation, and its contribution measurement rules are bound to insurance business scenarios. The interpretability enhancement module is used to generate the logical basis and visual explanation results for model decisions, and supports the output of explanations for insurance regulatory compliance. The model optimization module is used to dynamically adjust the model structure and training strategy to improve the model's generalization ability. It also has a built-in insurance business cycle adaptation unit, which is used to dynamically optimize the model update frequency according to business characteristics. Business effects include insurance period, claims processing time, etc. The risk assessment module is used to output the credit risk level and assessment results of insurance customers based on the global model, and supports the linked output of risk level and insurance product pricing rules.
[0021] It should be noted that the data preprocessing module communicates with the federated training module, the federated training module communicates bidirectionally with the privacy protection module, the model optimization module, and the risk assessment module, the blockchain incentive module communicates unidirectionally with the federated training module and the data preprocessing module, and the interpretability enhancement module communicates bidirectionally with the federated training module and the risk assessment module.
[0022] Specifically, the data preprocessing module includes a secure ID alignment unit, a multimodal feature unified representation unit, a feature standardization unit, and an insurance industry-specific risk feature extraction unit; The secure ID alignment unit employs privacy set intersection technology and, based on the KKRT protocol, achieves de-identification matching of cross-institutional customer identities, specifically as follows: Identify any two participating institutions, denoted as A and B, and set the customer IDs of participating institution A as follows: The set of customer IDs of participating institution B is denoted as . Through the combined operation of oblivious transfer and hash encryption, the core formula for calculating the intersection is: ; In the formula, This represents the privacy set intersection algorithm, which outputs a set of de-identified customer IDs shared across organizations. Simultaneously, it supports multi-dimensional cross-validation of insurance customer unique identifiers (such as policy number hashes and ID card hashes) through formulas. ( (Using different hash functions respectively) ensures the uniqueness of ID matching; through the secure ID alignment unit, it is possible to determine the customer samples jointly evaluated across institutions without disclosing non-intersecting customer IDs, providing a unified sample basis for subsequent federated learning, while avoiding ID matching errors through multi-dimensional hash verification.
[0023] The multimodal feature unified representation unit transforms textual, numerical, and time-series insurance credit-related data into feature vectors of a unified dimension through a multimodal large model, specifically including: For text-based data (such as claims descriptions and insurance declarations), semantic vectors are generated through a word embedding layer: ,in, Here, d represents the pre-trained word embedding function, d represents the embedding dimension, and Text represents text-based insurance credit data, such as unstructured text information like claims descriptions and insurance declarations. For numerical data (such as premium amount and payment frequency), standardize the data first and then map it into a vector: ,in, ( The mean, Standardization operations are represented by the standard deviation. For the linear mapping layer, xr represents numerical insurance credit data, such as premium amount, payment frequency and other numerical data; For time-series data (such as payment time series and claims history), a self-attention mechanism is used to capture time-series dependencies: ,in, For the self-attention calculation module, Sequence represents time-series insurance credit data, such as payment time sequence, claims history and other sequential data with a time relationship; Finally, features are aggregated through a multimodal fusion layer: ,in, Indicates feature splicing, For a multilayer perceptron, the output is a feature vector of uniform dimension d; By using a unified multimodal feature representation unit to transform insurance credit-related data into unified dimensional features, the modal differences between text, numerical, and time-series data are eliminated, enabling multi-source heterogeneous data to participate in federated training in the same feature space, thereby improving the model's ability to characterize the multi-dimensional risk features of insurance customers.
[0024] The feature standardization unit uses a local independent processing method to complete feature cleaning and normalization transformation: This unit uses Z-score normalization to perform local independent processing of features, as shown in the formula: ; Where x is the original feature value, The characteristic mean, The characteristic standard deviation, These are the standardized feature values, where i represents the index of the feature sample. Let represent the value of the i-th feature sample, and n represent the number of feature samples; all standardization operations are performed locally at the participating institutions to avoid the transfer of raw data across institutions. By standardizing features into a single unit, features of different dimensions and distributions are normalized to the same scale, eliminating the interference of dimensions on the training of federated learning models and ensuring the stability of model convergence and training efficiency.
[0025] The insurance industry-specific risk feature extraction unit automatically generates insurance industry-specific risk features based on actuarial rules. These features include the policy's continuous premium payment duration T, the volatility of claims amount V, and the risk coefficient of the insurance product combination R. The calculation logic for each risk characteristic is as follows: Continuous premium payment period for the policy: ,in The timestamp of the customer's i1th payment is used. i1 iterates through all the customer's payment records and calculates the duration of continuous payments by taking the difference between the maximum and minimum values of these timestamps. Volatility of claim amount ,in This indicates the amount of the customer's i2nd claim. It is a function of standard deviation. It is a mean function, and measures the volatility risk of claim amounts by calculating the ratio of the standard deviation of these claim amounts to the mean. Risk coefficient of insurance combination ,in The weight of the j-th insurance type (set by the actuary based on risk exposure). For the number of insurance types, is the basic risk coefficient for the j-th insurance type, used to quantify the comprehensive risk under a combination of multiple insurance types; By extracting risk features specific to the insurance industry, we can uncover risk factors unique to insurance business scenarios, such as payment continuity, claims volatility, and risks associated with insurance product combinations. This will compensate for the shortcomings of general features in characterizing insurance customer risks and significantly improve the industry-specific relevance of credit risk assessment.
[0026] Specifically, the federated training module supports adaptive switching between vertical federated learning and horizontal federated learning, and includes hierarchical training units based on insurance risk levels. Vertical federated learning is suitable for scenarios with overlapping samples and complementary features across institutions (such as collaboration between insurance companies and medical institutions, where the samples are from the same group of customers but hold insurance features and medical features respectively). It achieves collaborative computation of cross-institutional cross-features through secure multi-party computation protocols (such as a two-party secure logistic regression protocol). When using a logistic regression model in the vertical federated learning units, the objective function for joint optimization is: Where N represents the number of cross-institutional clients between participating institutions A and B. This is indicated as a customer risk label. Let i represent the feature of the i-th sample from participating institution A. Let represent the feature of the i-th sample from participating institution B. These represent the model weight vectors corresponding to participating institutions A and B, respectively. Let represent the transpose of the model weight vectors corresponding to participating institutions A and B, respectively, and let b represent the model bias term. This represents the Sigmoid activation function, which maps the output to the (0,1) interval to represent the risk probability; Within a secure multi-party computation framework, participating organizations A and B each compute gradients locally (e.g., for...). The gradient is updated through encrypted transmission and collaborative computation, ensuring the original features are preserved throughout the process. , It does not disclose data to external parties; by supporting vertical federated learning, it integrates complementary features across institutions (such as "insurance premium stability + medical chronic disease history") under the premise of strictly protecting data privacy, and significantly improves the assessment accuracy of scenarios such as health insurance underwriting and critical illness insurance risk pricing.
[0027] Horizontal federated learning is suitable for scenarios where cross-institutional features are consistent and samples are complementary (e.g., multiple property insurance companies collect features such as "number of auto insurance claims and accident liability ratio," but the samples cover different regions). An improved federated averaging algorithm is used to aggregate model parameters, and the formula is as follows: ;in Let M represent the global model parameters in the (k+1)th round, and M represent the total number of institutions participating in the federated training. This represents the number of local samples from participating institution j. This represents the model AUC value on the local validation set of participating institution j, used to correct for aggregation bias caused by differences in data quality. This represents the model parameters of participating institution j after training E rounds locally. By introducing a weighting mechanism for data quality and performance, institutions with high AUC values and excellent data quality are given higher aggregation weights to alleviate model bias caused by non-independent and identically distributed data. Through horizontal federated learning, the model integrates scattered samples from multiple institutions, enabling it to learn risk patterns that are common across regions (such as the fraud characteristics of "insurance fraud in other locations + multiple claims"), thereby improving the coverage and accuracy of anti-fraud and differentiated pricing of auto insurance premiums.
[0028] The stratified training unit based on insurance risk levels stratifies customer samples according to preset risk levels, including low risk. Medium risk and high risk For high-risk samples, a weighted loss and reinforcement iteration strategy is adopted, and the modified loss function is as follows: ,in Let represent the weighted loss function, which represents the total loss after considering the insurance risk level weights. It is used to measure the difference between the model's prediction and the actual situation. Q represents the number of insurance customer samples participating in the training, and h represents the customer sample index, used to traverse each customer sample. This represents the risk level weighting function. This indicates the risk level of the h-th customer. This means assigning different weights to samples based on their risk level. Cross-entropy loss function is a commonly used loss calculation method in classification tasks. This represents the true risk label of the h-th customer. This represents the model's predicted risk result for the h-th customer. Used to measure predicted values With real labels Differences; At the same time, the number of local training iterations is set for high-risk layer samples. , The preset number of iterations is used; through stratified training units based on insurance risk levels, the model enhances its sensitivity to risk identification of high-risk customers (such as those with frequent large claims or speculative policyholders with high insured amounts) by weight enhancement and iterative reinforcement, thereby reducing the risk of huge payouts due to missed detections.
[0029] The federated training module also includes a large model-assisted pre-training unit and a global alignment unit. The large model-assisted pre-training unit utilizes a server-side multimodal large language model for global pre-training. The pre-training tasks include: Masked feature prediction uses random masked multimodal features to allow the model to reconstruct the original features. The loss function is... ,in This indicates a feature masking operation. For mean square error loss, Represents the multimodal feature vector of an insurance customer; Risk label semantic alignment semantically associates features with predefined risk labels, using the following loss function: Where K is the number of risk categories, Let be the one-hot label of class e for the i-th sample; the large model-assisted pre-training unit uses the general feature extraction and semantic understanding capabilities of the large model to learn the global feature representation of insurance data in unsupervised / weakly supervised scenarios, providing a high-quality initialization model for subsequent federated training, accelerating convergence and improving generalization performance; Global alignment units are based on pre-trained models Global semantic alignment is performed on the locally fine-tuned model using the KL divergence loss function: ,in KL divergence is used to measure the local model. With global pre-trained models The difference in predicted distribution is minimized to ensure the consistency of local models in the global semantic space. The global alignment unit solves the "semantic drift" problem of local models caused by the difference in data distribution among institutions, ensuring the consistency of global model decisions in cross-institutional scenarios after federated training, and meeting the insurance industry's requirements for the uniformity of risk assessment standards.
[0030] Specifically, the federated training module employs a dynamic adaptive aggregation algorithm. Based on the data distribution characteristics, data quality, and model performance of each participating institution, it dynamically adjusts the aggregation weights of the participating parameters. Furthermore, it assigns weighting coefficients to institutions containing high-value risk characteristics to mitigate the model bias caused by non-independent and identically distributed data. Specifically: Multi-dimensional indicator evaluation steps: Through the built-in statistical analysis component, evaluate the data distribution characteristics, data quality and local model performance of each participating institution; The statistical analysis component, as an existing technology, evaluates the target dimensions of each participating institution through the following specific indicators to ensure that the evaluation process is open, transparent, and feasible: When assessing the data distribution characteristics, the risk level distribution deviation (which measures the deviation of the proportion of institutional clients’ risk levels from the industry average), the geographical coverage overlap rate (which measures the degree of overlap between the geographical coverage of institutional clients and the geographical profile of the overall insurance clients) and the preset weights are used to calculate the results. When assessing data quality, the feature missing rate (the proportion of missing values in the feature data of a statistical institution), the data update cycle (the frequency of data updates of the recording institution), and a preset weight are used to calculate the data. When evaluating the performance of the local model, the validation set AUC value (which measures the discriminative ability of the local model in risk prediction), the risk prediction accuracy (which is the ratio of the local model's prediction results to the actual risk labels) and the preset weights are used to calculate the results. Preliminary weight allocation steps: Based on the above evaluation results, including data distribution characteristics, data quality, and local model performance, preliminary aggregate weights for each institution are generated using a weighted summation method; High-value characteristic weighting step: For participating institutions holding high-value risk characteristics, a weighting coefficient with a set multiple is assigned on the basis of the initial weight to obtain the final determined weight; Global parameter aggregation steps: According to the final determined weights, the encrypted local model parameters uploaded by each institution are weighted and summed to generate global credit risk assessment model parameters. Through additional weighting, the parameters of institutions with these characteristics can be given a higher proportion in the global aggregation, ensuring that the key risk information they carry is fully integrated into the global model. This effectively alleviates the model bias caused by the uneven distribution of data among institutions, enabling the global model to capture various credit risk patterns of insurance customers more comprehensively and accurately.
[0031] Specifically, the privacy protection module adopts a hybrid privacy protection mechanism that combines homomorphic encryption, differential privacy, and trusted execution environment, and includes a risk data sensitivity classification unit; The specific configuration of the risk data sensitivity classification unit is as follows: Insurance data is divided into core sensitive data, general sensitive data, and non-sensitive data. Core sensitive data includes insurance customers' medical history and asset certificates, general sensitive data includes customers' insurance records and payment amounts, and non-sensitive data includes publicly available insurance information. Corresponding encryption strengths are configured: core sensitive data uses the highest encryption strength, general sensitive data uses medium encryption strength, and non-sensitive data uses simplified encryption strength. The specific application methods of the hybrid privacy protection mechanism are as follows: Homomorphic encryption is used to perform cross-institutional feature calculations and model parameter aggregation in a encrypted state, enabling joint operations to be carried out without decryption; Differential privacy is used to add noise perturbations to the model gradients or parameters uploaded by each participating institution, and to adjust the noise adaptation strength for core sensitive data and general sensitive data respectively. Trusted Execution Environment (TEE) constructs an independent and secure computing space at the hardware level, isolating sensitive data and computing logic during the training process within this space to prevent external access and tampering. By using a risk data sensitivity classification unit, the problem of excessive encryption affecting efficiency and insufficient encryption posing risks caused by adopting a uniform privacy protection strategy for data with different sensitivity levels in the insurance industry is solved. By configuring encryption strength differently according to data sensitivity, the encryption overhead of general and non-sensitive data is reduced while ensuring the security of core sensitive data, thereby improving the adaptability and execution efficiency of the privacy protection mechanism.
[0032] Specifically, the blockchain incentive module is implemented based on a consortium blockchain architecture, including a contribution measurement unit and a smart contract incentive unit; The specific methods by which the contribution metric unit measures each dimension are as follows: The data volume is measured by the number of valid samples provided by participating institutions; the data quality is measured by the feature missing rate and the correlation with insurance risks; the computing resource investment is measured by the local model training time and the scale of computing power investment; and the model performance contribution is measured by the improvement of the AUC value of the local model on the global model validation set. For data that provides key anti-fraud features (including historical insurance fraud records, abnormal insurance behavior patterns, and cross-institutional fraud correlation data), an additional weight of a set multiple is given on the basis of the comprehensive score of the above dimensions. The specific configuration of the smart contract incentive unit is as follows: The incentive allocation rules include the share of model usage rights (allocating global model call permissions according to the contribution score ratio) and the premium sharing ratio (associating insurance business revenue sharing according to the contribution score); the generated on-chain evidence records include the participating institution's identifier, contribution score, incentive type and amount, and allocation timestamp; it supports the connection with the incentive settlement interface of the insurance core business system and financial settlement system to realize automatic reconciliation and transfer of incentive amounts.
[0033] Specifically, the interpretability enhancement module includes a local decision interpretation unit, a global aggregation visualization unit, and a regulatory compliance interpretation unit; The local decision interpretation unit uses the CART decision tree model to construct a logical tree linking "input features - risk assessment results". By traversing the decision path, it matches the mapping relationship between the feature combination of a single customer and the risk score, and outputs key influencing factors, including but not limited to feature conditions directly related to risk level upgrades such as "more than 5 claims in the past 3 years", "two consecutive overdue payments", and "claims filed within 1 month after purchasing a high-coverage insurance policy". The specific configuration of the global aggregation visualization unit is as follows: the contribution of each participating institution's input features is ranked through feature importance analysis, principal component analysis is used to reduce the dimensionality of high-dimensional features, and the generated visualization view includes a heatmap of the weight distribution of each institution's features in the global model, a curve showing the change between the federated aggregation rounds and the model's AUC value, and a pie chart showing the contribution of core features to the risk assessment results. The regulatory compliance interpretation unit is specifically configured as follows: Based on the regulatory requirements of the China Banking and Insurance Regulatory Commission (CBIRC) such as the "Insurance Technology Supervision Measures" and the "Personal Insurance Information Disclosure Management Measures", it automatically generates structured reports. The reports include an explanation of the federated learning training framework, core risk characteristic screening criteria (such as the correlation threshold between characteristics and risk labels), and the basis for setting thresholds for risk level classification (such as the probability intervals corresponding to low / medium / high risk). It supports exporting the reports in PDF format, which can be directly used for regulatory inspections and public disclosure of customer risk assessment results.
[0034] Specifically, the model optimization module includes an incremental learning unit, a model compression unit, an edge computing adaptation unit, and an insurance business cycle adaptation unit; The incremental learning unit is specifically configured as follows: Supported new data includes monthly insurance claims data, quarterly customer credit update data, and new insurance business data. When the new data accumulates to a preset threshold, each participating institution only needs to calculate and upload the incremental values of the model parameters corresponding to the new data. After receiving the data, the federated training module performs weighted aggregation with the global model parameters to complete the dynamic update of the global model without having to start the full federated training process. The specific configuration of the model compression unit is as follows: Knowledge distillation uses the trained global credit risk assessment model as the teacher model and the local lightweight models of each institution as the student model. By minimizing the difference in prediction distribution between the teacher and student models, the core risk feature extraction capability is distilled. Parameter quantization converts the model's 32-bit floating-point parameters into 8-bit integer parameters, reducing model storage and cross-institutional transmission overhead while controlling the model performance loss to no more than 5%. The specific configuration of the edge computing adapter unit is as follows: For edge nodes such as local underwriting terminals and branch business servers of insurance companies, redundant feature processing layers and computing units in the global model are removed, and the model inference logic is optimized to adapt to the computing power of edge devices. After the adaptation is completed, the edge nodes can directly load the optimized model and realize credit risk assessment based on local customer feature data without relying on the federated center server. The specific configuration of the insurance business cycle adaptation unit is as follows: For short-term insurance products such as one-year short-term insurance and six-month accident insurance, the model update frequency is set to once a month; for long-term products such as five-year long-term life insurance and ten-year annuity insurance, the model update frequency is set to once a quarter; for fast-claim insurance products, an additional incremental update is added every two weeks based on the characteristics of claim processing time, to ensure that the model adapts to the business rhythm. This model optimization module achieves dynamic global model updates through incremental learning, reduces resource consumption through model compression, enables rapid local evaluation through edge adaptation, and adjusts the update frequency according to the insurance business cycle. Ultimately, it achieves comprehensive optimization of the model in terms of efficiency, resource consumption, and scenario adaptability, meeting the dynamic and lightweight requirements of insurance credit risk assessment.
[0035] This invention also provides a method for cross-institutional credit risk assessment of insurance clients based on the above system, the method comprising the following steps: Step 1: Cross-institutional participants complete the local data security ID alignment, multimodal feature unified representation, standardization processing, and insurance industry-specific risk feature extraction through the data preprocessing module to generate a local feature dataset; Step 2: The federated training module receives training requests from each participant, adaptively selects vertical or horizontal federated learning mode based on data distribution characteristics, and conducts local model training and parameter updates in layers according to insurance risk level based on the risk data sensitivity classification protection mechanism of the privacy protection module. Before training, global pre-training is completed through the server-side multimodal large language model, and then the pre-trained model is distributed to each participant for local fine-tuning. After fine-tuning, semantic alignment is achieved through the global alignment unit. Step 3: The federated training module uses a dynamic adaptive aggregation algorithm to globally aggregate the encrypted model parameters uploaded by each participant, assign weighting coefficients to high-value risk characteristic institutions, and generate an initial global model. Step 4: The interpretability enhancement module performs decision logic analysis, visualization processing, and regulatory compliance report generation on the initial global model, outputting a model interpretability report; Step 5: The model optimization module, in conjunction with the characteristics of the insurance business cycle, performs incremental updates, compression, and edge adaptation optimization on the initial global model to generate the final global credit risk assessment model. Step Six: Based on the final global model, the risk assessment module classifies the credit risk of the target insurance customers and outputs assessment results that match the pricing rules of the insurance products. Step 7: The blockchain incentive module automatically completes the incentive distribution and on-chain storage based on the contributions of each participant during the training process through smart contracts.
[0036] All formulas in this solution are based on dimensionless numerical calculations. Dimensionlessness can be achieved through conventional methods such as standardization, and the specific process will not be elaborated here. The formulas are obtained by software simulation and fitting of a large amount of data, which can approximate the real scenario. The preset parameters in the formulas are configured by those skilled in the art according to the actual business scenario.
[0037] This embodiment can be implemented by software, hardware, firmware, or any combination thereof; if implemented by software, it can be in the form of a computer program product, and after the computer instructions are loaded or executed, the process and function described in this solution can be realized. The computer instructions can be transmitted between different storage media via wired or wireless means.
[0038] The program numbers in this solution do not represent the execution order. The specific execution order is determined by the functional logic and does not constitute an implementation restriction. Each example unit and algorithm step can be implemented by electronic hardware or a combination of hardware and software. The specific implementation method depends on the application scenario and design constraints.
[0039] In this solution, the unit division is only a logical functional division. The actual implementation can be adjusted as needed, and multiple units can be integrated or split. If the functional unit is implemented in software and sold / used independently, it can be stored in a computer-readable storage medium (including USB flash drive, ROM, RAM, magnetic disk, optical disk, etc.), and the relevant instructions can drive the computer device to execute the corresponding method steps.
[0040] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0041] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A cross-institutional credit risk assessment system for insurance clients based on federated learning, characterized in that: include: The data preprocessing module is used for secure alignment and standardization of multi-source data across institutions, generating feature datasets usable for federated training. It also has a built-in risk feature extraction unit specifically for the insurance industry, which is used to automatically generate insurance industry-specific risk features for policy data, claims records, and credit information. The federated training module is used to conduct distributed model training based on multi-party feature datasets. It generates a global credit risk assessment model through dynamic aggregation strategies and supports a hierarchical training mechanism based on insurance risk levels. For high-risk customer samples, a reinforcement learning strategy is used to optimize the model convergence direction. The privacy protection module is used to protect the privacy of data, model parameters and intermediate results during training, and has a built-in risk data sensitivity classification unit to dynamically adjust the privacy protection strength according to the sensitivity level of insurance data. The blockchain incentive module is used to quantify the contribution of cross-institutional data and computing resources, realize automated incentive allocation, and its contribution measurement rules are bound to insurance business scenarios. The interpretability enhancement module is used to generate the logical basis and visual explanation results for model decisions, and supports the output of explanations for insurance regulatory compliance. The model optimization module is used to dynamically adjust the model structure and training strategy to improve the model's generalization ability, and it also has a built-in insurance business cycle adaptation unit to dynamically optimize the model update frequency according to business characteristics. The risk assessment module is used to output the credit risk level and assessment results of insurance customers based on the global model, and supports the linked output of risk level and insurance product pricing rules.
2. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The data preprocessing module includes a secure ID alignment unit, a multimodal feature unified representation unit, a feature standardization unit, and an insurance industry-specific risk feature extraction unit. The secure ID alignment unit uses privacy set intersection technology to achieve de-identification matching of cross-institutional customer identities and supports multi-dimensional cross-verification of insurance customer unique identifiers. The multimodal feature unified representation unit transforms textual, numerical, and time-series insurance credit-related data into feature vectors of a unified dimension through a multimodal large model; The feature standardization unit uses a local independent processing method to complete feature cleaning and normalization transformation; The insurance industry-specific risk feature extraction unit automatically generates insurance industry-specific risk features based on actuarial rules.
3. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The federated training module supports adaptive switching between vertical federated learning and horizontal federated learning, and includes hierarchical training units based on insurance risk levels. The vertical federated learning is suitable for scenarios with overlapping samples and complementary features across institutions, and generates cross-institutional cross features through secure multi-party computation; The horizontal federated learning is suitable for scenarios where cross-institutional features are consistent and samples are complementary, and an improved federated averaging algorithm is used to aggregate model parameters. The stratified training unit based on insurance risk level stratifies customer samples according to preset risk levels and uses a higher number of iterations for reinforcement training of high-risk samples. The federated training module also includes a large model-assisted pre-training unit and a global alignment unit. The large model-assisted pre-training unit uses a server-side multimodal large language model for global pre-training, and the global alignment unit performs global semantic alignment on the locally fine-tuned model based on the pre-trained model.
4. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 3, characterized in that, The federated training module employs a dynamic adaptive aggregation algorithm, which dynamically adjusts the aggregation weights of the participating institutions' parameters based on the data distribution characteristics, data quality, and model performance of each participating institution, and assigns weighting coefficients to institutions containing high-value risk characteristics.
5. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The privacy protection module adopts a hybrid privacy protection mechanism that combines homomorphic encryption, differential privacy, and trusted execution environment, and includes a risk data sensitivity classification unit. The risk data sensitivity classification unit divides insurance data into core sensitive data, general sensitive data, and non-sensitive data, and configures different encryption strengths accordingly. The homomorphic encryption is used to support feature calculation and parameter aggregation in the dense state; The differential privacy is used to add noise perturbation to the uploaded model gradient or parameters, and the noise intensity is dynamically adjusted according to the data sensitivity. The trusted execution environment is used to construct a hardware-level secure computing space, isolating sensitive data and computational logic during the training process.
6. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The blockchain incentive module is implemented based on a consortium blockchain architecture and includes a contribution quantification unit and a smart contract incentive unit. The contribution measurement unit comprehensively scores the data provided by participating institutions based on the amount of data, data quality, computing resource input, and model performance contribution, and assigns extra weight to data that provides key anti-fraud features. The smart contract incentive unit automatically executes the incentive allocation rules based on the contribution score, generates an immutable incentive record and stores it on the blockchain, and supports integration with the incentive settlement interface of the insurance business system.
7. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The interpretability enhancement module includes a local decision interpretation unit, a global aggregation visualization unit, and a regulatory compliance interpretation unit; The local decision interpretation unit uses a decision tree model to match the correlation logic between credit risk assessment results and input features, and outputs the key influencing factors of a single customer's risk score; The global aggregation visualization unit generates a visual view of the model aggregation process through feature importance analysis and dimensionality reduction techniques. The regulatory compliance interpretation unit automatically generates a structured report containing model principles, feature selection criteria, and risk level classification standards in accordance with insurance regulatory requirements, thus meeting the disclosure requirements of insurance regulatory agencies.
8. The cross-institutional credit risk assessment system for insurance clients based on federated learning according to claim 1, characterized in that, The model optimization module includes an incremental learning unit, a model compression unit, an edge computing adaptation unit, and an insurance business cycle adaptation unit. The incremental learning unit supports dynamic model updates based on new data, without the need to retrain the global model. The model compression unit reduces model storage and transmission overhead through knowledge distillation and parameter quantization techniques; The edge computing adaptation unit is used to adapt the global model to the edge node to achieve rapid local risk assessment; The insurance business cycle adaptation unit dynamically adjusts the model update frequency based on the characteristics of the insurance product cycle and claims processing time.
9. A method for cross-institutional credit risk assessment of insurance clients based on the system described in any one of claims 1-8, characterized in that, The method includes the following steps: Step 1: Cross-institutional participants complete the local data security ID alignment, multimodal feature unified representation, standardization processing, and insurance industry-specific risk feature extraction through the data preprocessing module to generate a local feature dataset; Step 2: The federated training module receives training requests from each participant, adaptively selects vertical or horizontal federated learning mode based on data distribution characteristics, and conducts local model training and parameter updates in layers according to insurance risk level based on the risk data sensitivity classification protection mechanism of the privacy protection module. Before training, global pre-training is completed through the server-side multimodal large language model, and then the pre-trained model is distributed to each participant for local fine-tuning. After fine-tuning, semantic alignment is achieved through the global alignment unit. Step 3: The federated training module uses a dynamic adaptive aggregation algorithm to globally aggregate the encrypted model parameters uploaded by each participant, assign weighting coefficients to high-value risk characteristic institutions, and generate an initial global model. Step 4: The interpretability enhancement module performs decision logic analysis, visualization processing, and regulatory compliance report generation on the initial global model, outputting a model interpretability report; Step 5: The model optimization module, in conjunction with the characteristics of the insurance business cycle, performs incremental updates, compression, and edge adaptation optimization on the initial global model to generate the final global credit risk assessment model. Step Six: Based on the final global model, the risk assessment module classifies the credit risk of the target insurance customers and outputs assessment results that match the pricing rules of the insurance products. Step 7: The blockchain incentive module automatically completes the incentive distribution and on-chain storage based on the contributions of each participant during the training process through smart contracts.
Citation Information
Cited By
Small and micro enterprise joint risk portrait and abnormal behavior identification system
CN121961261A