Real-time underwriting risk control method, system and equipment based on containerized deployment and medium

By containerizing the underwriting risk control model jointly trained with XGBoost and neural networks, the problems of inflexible model deployment and insufficient high-concurrency processing capabilities in existing technologies are solved, realizing an efficient and stable real-time underwriting risk control system that can quickly adjust resource allocation and respond to user requests according to business needs.

CN120975927APending Publication Date: 2025-11-18TIANJIN HEALTH CARE BIG DATA CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510834878.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing underwriting risk control methods lack flexibility and high-concurrency processing capabilities in model deployment, real-time calculation, and interface implementation, making it impossible to respond to user requests in a timely manner and lacking dynamic optimization mechanisms, resulting in low efficiency in insurance business.

Method used

The underwriting risk control model is trained using XGBoost and neural networks, and then packaged into a Docker image along with the preprocessing module and FastAPI interface service. It is then deployed in a containerized manner through a Kubernetes cluster. Combined with Prometheus and Kubernetes HPA, dynamic resource adjustment is achieved to ensure the stable operation of the system in a high-concurrency environment.

Benefits of technology

It enables efficient deployment and elastic scaling of the model, improves the system's concurrent processing capabilities and response speed, ensures the accuracy and stability of real-time underwriting risk control, and dynamically adjusts underwriting strategies to adapt to business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975927A_ABST
    Figure CN120975927A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of underwriting risk control, and particularly relates to a real-time underwriting risk control method, system, equipment and medium based on containerized deployment, and the method comprises the steps: packaging a trained underwriting risk control model, a preprocessing module and an interface service script into a Docker mirror image; the method comprises the following steps: deploying a Docker mirror image as a container, and configuring: when each container is started, a FastAPI service is automatically bound to a port; mapping a container port into a cluster unified entry; an external request is distributed to the container through load balancing and routed to a corresponding FastAPI interface, the preprocessing module is triggered to analyze data, and an underwriting risk control model is called to calculate a risk score; the returned risk score dynamically adjusts the underwriting strategy; when the risk score exceeds a threshold value, an underwriting rejection decision is returned; and when the risk score is lower than a threshold value, a low insurance premium or standard insurance premium suggestion is returned. The concurrent processing capability and the response speed of the system are improved, and the stable operation of the real-time underwriting risk control service is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of underwriting risk control technology, specifically relating to a real-time underwriting risk control method, system, equipment, and medium based on containerized deployment. Background Technology

[0002] Insurance underwriting is the process by which insurance companies review, verify, and select risks in an insured's application for insurance. It is a prerequisite for insurance underwriting, a key link in the insurer's business operations, and an important condition for ensuring the stable operation of insurance companies.

[0003] With the rapid development of artificial intelligence technology, machine learning models are widely used in the insurance industry, playing a crucial role, especially in risk control assessment during the underwriting process. However, existing underwriting risk control methods have several shortcomings in model deployment, real-time computation, and interface implementation. On the one hand, model deployment lacks flexibility and scalability, making it difficult to quickly adjust resource allocation according to actual business needs and efficiently cope with dynamic changes in business traffic. On the other hand, existing methods cannot provide timely responses to high-concurrency requests, resulting in a poor user experience and limiting the processing efficiency of insurance business. Furthermore, existing methods lack targeted data monitoring and dynamic optimization mechanisms, making it impossible to adjust underwriting strategies in a timely manner based on real-time data feedback, thus hindering accurate risk assessment and decision-making.

[0004] Therefore, how to efficiently deploy underwriting risk control models and build a real-time underwriting risk control system with high concurrency processing capabilities, real-time response, and dynamic optimization mechanisms has become a technical problem that the insurance industry urgently needs to solve. Summary of the Invention

[0005] In view of the above-mentioned shortcomings of the prior art, the present invention provides a real-time underwriting risk control method, system, equipment and medium based on containerized deployment to solve the above-mentioned technical problems.

[0006] In a first aspect, the technical solution of the present invention provides a real-time underwriting risk control method based on containerized deployment, including: S1. Based on historical underwriting data, an underwriting risk control model is jointly trained using XGBoost and neural networks. The underwriting risk control model is used to calculate a risk score; the risk score is used to quantify the underwriting risk level. S2. Package the trained underwriting risk control model, preprocessing module, and FastAPI interface service script into a Docker image; S3. Deploy Docker images as multi-instance services (i.e., containers) through a Kubernetes cluster and configure them as follows: When each container starts, the FastAPI service automatically binds to a specified port within the container and listens on that port; The container port is mapped to a unified entry point for the cluster to receive external requests through the Service resource; S4. External requests are distributed to containers via Service load balancing. The port bound to the container receives the requests and routes them to the corresponding FastAPI interface, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score; the calculation result is returned through the original port path. S5. Dynamically adjust the underwriting strategy based on real-time data feedback and returned risk scores: If the risk score exceeds the threshold, a decision to refuse coverage will be made. If the risk score is below the threshold, a low premium or standard premium recommendation will be given.

[0007] By employing XGBoost and neural networks to jointly train the underwriting risk control model, the advantages of both models can be fully leveraged to improve the accuracy of risk scoring, thereby more accurately quantifying the underwriting risk level. The model, preprocessing module, and FastAPI interface service scripts are encapsulated as Docker images and deployed as multi-instance services using a Kubernetes cluster, achieving efficient deployment and elastic scaling of the model. This deployment method allows for flexible adjustment of the number of containers according to business needs, effectively handling high-concurrency requests and ensuring the system's real-time response capability. It solves the efficiency bottlenecks of existing underwriting risk control methods in deployment, real-time computation, and interface implementation, providing the insurance industry with an efficient and reliable real-time underwriting risk control solution.

[0008] As a further limitation of the technical solution of the present invention, the method also includes: Prometheus collects CPU utilization and FastAPI response time metrics for each container in real time. When any metric exceeds the corresponding preset threshold and remains so for a set time, Kubernetes HPA is triggered to automatically increase the number of containers. When all metrics are below half of the set threshold and remain so for a set time, Kubernetes HPA is triggered to automatically decrease the number of containers. in,

[0009]

[0010] In the formula, n is the total number of sampling requests.

[0011] By collecting real-time metrics such as CPU utilization and FastAPI response time for each container, and automatically adjusting the number of containers based on these metrics, the system can promptly add containers when business traffic increases to ensure system processing capacity, and automatically reduce containers when business traffic decreases to avoid resource waste. This dynamic resource adjustment mechanism further improves the system's resource utilization and responsiveness, ensuring that the system maintains efficient and stable operation under different business loads, providing strong support for real-time underwriting and risk control.

[0012] As a further limitation of the technical solution of the present invention, step S1, which involves jointly training the underwriting risk control model using XGBoost and a neural network, includes: Preprocess historical underwriting data, customer information, and risk characteristics; The XGBoost model is used to train the preprocessed data. During the training process, the XGBoost model automatically calculates the importance score of each feature. The importance score of each feature represents the degree of contribution of the feature to the model. All features are sorted according to their importance scores, and features with importance scores higher than a preset threshold are selected as key features. A neural network model is constructed based on the selected key features; The output features of the XGBoost model and the preprocessed original features are fed into the neural network model for training. The model parameters are adjusted using the Adam algorithm to minimize the loss function. The prediction results of the XGBoost model and the neural network model are fused to obtain the final underwriting risk control model prediction result; when the prediction result meets the set requirements, the model training is completed.

[0013] By preprocessing historical underwriting data, customer information, and risk characteristics, and using the XGBoost model to calculate feature importance scores, key features are selected. This reduces the impact of redundant information on the model, improving training efficiency and accuracy. A neural network model is then constructed based on these selected key features. The output features of the XGBoost model are input together with the preprocessed original features into the neural network model for training. By fusing the prediction results of the two models, the model's performance is further improved, making risk scoring more accurate and reliable, and providing a more scientific basis for underwriting decisions.

[0014] As a further limitation of the technical solution of the present invention, the steps for calculating the importance score of each feature during the XGBoost model training process include: During the training of each tree, for each split point of each feature, the gain brought by the split is calculated; For each feature, the gains of that feature in all trees are summed to obtain the total gain of the feature; The importance score of a feature is obtained by dividing its total gain by the sum of the total gains of all features. The specific calculation formula is as follows:

[0015] In the formula, s represents the total number of trees in the model. m This represents the total number of features.

[0016] By calculating the gain from each feature split during the training of each tree, and then summing and normalizing the gains of that feature across all trees, an importance score for the feature is obtained. This method objectively evaluates the contribution of each feature to the model, providing an accurate basis for feature selection, and helps optimize the model structure, improve the model's generalization ability, and predictive accuracy.

[0017] As a further limitation of the technical solution of the present invention, the gain calculation method includes: for each split point of each feature, calculating the Gini impurity before and after splitting, and calculating the gain based on the Gini impurity before and after splitting.

[0018]

[0019] In the formula, Indicates the first The proportion of samples of each class in the dataset, where Q represents the total number of classes. For the split subset of data The number of samples, Data set before splitting The number of samples.

[0020] Gini impurity is a commonly used metric for measuring dataset purity. By calculating the change in Gini impurity before and after splitting, we can accurately assess the extent to which feature splitting improves model performance, thereby more reasonably calculating feature importance scores, further optimizing the model training process, and improving the model's prediction performance.

[0021] As a further limitation of the technical solution of the present invention, step S4 specifically includes: Once an external request arrives at the Kubernetes Service, it is load-balanced and distributed to the containers. The container receives requests through the port bound to the FastAPI service, and FastAPI routes them to the corresponding interface processing function. The interface processing function parses the request to obtain the required data, triggers the preprocessing module in the container to parse the data, and calls the underwriting risk control model to calculate the risk score. The calculation results will be returned via the original port path.

[0022] Kubernetes Service's load balancing functionality distributes requests to various containers. Containers receive requests through the port bound to the FastAPI service, which then routes them to the corresponding API processing functions. This approach ensures that requests are processed efficiently and accurately, improving the system's concurrency and response speed, and guaranteeing the stable operation of the real-time underwriting and risk control service.

[0023] As a further limitation of the technical solution of the present invention, the steps of parsing the request to obtain the required data, triggering the preprocessing module in the container to parse the data, and calling the underwriting risk control model to calculate the risk score include: The interface processing function extracts the customer's basic information, historical claims records, and health status characteristics from the request; it then triggers the preprocessing module within the container to preprocess the extracted features to make them conform to the format requirements of the model input. The extracted features are input into the trained XGBoost model to obtain the intermediate feature representation output by the XGBoost model; at the same time, the preprocessed original features are combined with the intermediate features output by XGBoost to form an enhanced feature set. The enhanced feature set is input into the trained neural network model, and the risk score for each customer is calculated through the forward propagation of the neural network model. The calculated risk scores are then standardized. Among them, risk score

[0024] In the formula, For the first Each feature weight, For the first There are 1 feature values, where m is the total number of features.

[0025] By extracting basic customer information, historical claims records, and health status characteristics from the request and preprocessing them to conform to the model input format requirements, the extracted features are then input into the XGBoost model to obtain intermediate feature representations. These intermediate representations are combined with the preprocessed original features to form an enhanced feature set, which is then input into a neural network model for risk scoring. The calculation results are then standardized. This series of steps fully utilizes various feature information, improving the accuracy and stability of risk scoring and providing more reliable support for underwriting decisions.

[0026] Secondly, the technical solution of the present invention also provides a real-time underwriting risk control system based on containerized deployment, comprising: The model training module, based on historical underwriting data, uses XGBoost and neural networks to jointly train the underwriting risk control model. The underwriting risk control model is used to calculate the risk score; the risk score is used to quantify the underwriting risk level. The containerization deployment module is used to encapsulate the trained underwriting and risk control model, preprocessing module, and FastAPI interface service script into a Docker image; deploy the Docker image as a multi-instance service, i.e., a container, through a Kubernetes cluster, and configure it so that when each container starts, the FastAPI service automatically binds to the specified port inside the container and listens on the port; and map the container port to a unified cluster entry point to receive external requests through the Service resource. The risk control interface engine distributes external requests to containers via Service load balancing. The containers receive requests on their bound ports and route them to the corresponding FastAPI interfaces, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score. The calculation results are then returned through the original port path. The dynamic decision-making module is used to dynamically adjust the underwriting strategy based on real-time data feedback and the returned risk score: when the risk score exceeds the threshold, a decision to refuse underwriting is returned; when the risk score is below the threshold, a suggestion of low premium or standard premium is returned.

[0027] The system comprises a model training module, a containerized deployment module, a risk control interface engine, and a dynamic decision-making module. The model training module improves the accuracy of risk scoring by jointly training XGBoost and neural network models. The containerized deployment module enables efficient model deployment and elastic scaling, enhancing the system's concurrent processing capabilities and response speed. The risk control interface engine ensures that external requests are processed promptly and accurately, achieving real-time underwriting risk control. The dynamic decision-making module dynamically adjusts underwriting strategies based on real-time data feedback and returned risk scores, achieving precise risk assessment and decision-making. All modules work collaboratively, resolving the efficiency bottlenecks in deployment, real-time computation, and interface implementation of existing underwriting risk control methods.

[0028] As a further limitation of the technical solution of the present invention, the containerized deployment module is also used to collect the CPU utilization and FastAPI interface response time of each container in real time through Prometheus; when any indicator exceeds the corresponding preset threshold and continues for a set time, Kubernetes HPA is triggered to automatically increase the number of containers; when all indicators are lower than half of the set threshold and continue for a preset time, Kubernetes HPA is triggered to automatically reduce the number of containers. in,

[0029]

[0030] In the formula, n is the total number of sampling requests.

[0031] As a further limitation of the technical solution of the present invention, the model training module includes a data preprocessing unit, a first-order training unit, a second-order training unit, and a training processing unit. The data preprocessing unit is used to preprocess historical underwriting data, customer information, and risk characteristics; The first-order training unit is used to train the preprocessed data using the XGBoost model. During the training process, the XGBoost model automatically calculates the importance score of each feature. The importance score of each feature is obtained, which represents the contribution of the feature to the model. All features are sorted according to their importance scores, and features with importance scores higher than a preset threshold are selected as key features. The second-order training unit is used to build a neural network model based on the selected key features. The output features of the XGBoost model and the preprocessed original features are input into the neural network model for training. The model parameters are adjusted by the Adam algorithm to minimize the loss function. The training processing unit is used to fuse the prediction results of the XGBoost model and the neural network model to obtain the final underwriting risk control model prediction result; when the prediction result meets the set requirements, the model training is completed.

[0032] As a further limitation of the technical solution of the present invention, the first-order training unit is also used to calculate the gain brought by splitting at each split point of each feature during the training process of each tree; for each feature, the gains of the feature in all trees are accumulated to obtain the total gain of the feature; the total gain of the feature is divided by the sum of the total gains of all features to obtain the importance score of the feature, and the specific calculation formula is as follows:

[0033] In the formula, s represents the total number of trees in the model. m This represents the total number of features.

[0034] As a further limitation of the technical solution of the present invention, the gain calculation method includes: for each split point of each feature, calculating the Gini impurity before and after splitting, and calculating the gain based on the Gini impurity before and after splitting.

[0035]

[0036] In the formula, Indicates the first The proportion of samples of each class in the dataset, where Q represents the total number of classes. For the split subset of data The number of samples, Data set before splitting The number of samples.

[0037] As a further limitation of the technical solution of the present invention, the risk control interface engine is specifically used to load balance and distribute external requests to each container after they arrive at the Kubernetes Service; the container receives the request through the port bound to the FastAPI service, and FastAPI routes it to the corresponding interface processing function; the interface processing function parses the request to obtain the required data, triggers the preprocessing module in the container to parse the data, calls the underwriting risk control model to perform risk scoring calculation, and returns the calculation result through the original port path.

[0038] As a further limitation of the technical solution of the present invention, the interface processing function extracts the customer's basic information, historical claims records, and health status characteristics from the request; and triggers the preprocessing module in the container to preprocess the extracted features to make them conform to the format requirements of the model input. The risk control interface engine inputs the extracted features into the trained XGBoost model to obtain the intermediate feature representation output by the XGBoost model; at the same time, it combines the preprocessed original features with the intermediate features output by XGBoost to form an enhanced feature set; the enhanced feature set is input into the trained neural network model, and the risk score for each customer is calculated through the forward propagation of the neural network model; the calculated risk score is then standardized. Among them, risk score

[0039] In the formula, For the first Each feature weight, For the first There are 1 feature values, where m is the total number of features.

[0040] Thirdly, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing computer program instructions executable by the at least one processor, the computer program instructions being executed by the at least one processor to enable the at least one processor to execute the real-time underwriting and risk control method based on containerized deployment as described in the first aspect.

[0041] Fourthly, the present invention also provides a non-transitory computer-readable storage medium that stores computer instructions that cause the computer to execute the real-time underwriting and risk control method based on containerized deployment as described in the first aspect.

[0042] The beneficial effects of this invention are that it enables rapid deployment of underwriting models and implementation of underwriting risk control interfaces, ensures the accuracy of model deployment, reduces redundant code, improves deployment and upgrade efficiency, and helps insurance companies to quickly and efficiently promote business development.

[0043] By using containerization technology and high-performance API frameworks (such as FastAPI), we have achieved efficient deployment and invocation of underwriting risk control models, enabling rapid response to user requests in high-concurrency environments and ensuring the real-time nature of insurance business. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic flowchart illustrating a method according to an embodiment of the present invention.

[0046] Figure 2 This is a schematic block diagram of a system according to an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the specific embodiments. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] like Figure 1 As shown, this embodiment of the invention provides a real-time underwriting risk control method based on containerized deployment, including: S1. Based on historical underwriting data, an underwriting risk control model is jointly trained using XGBoost and neural networks. The underwriting risk control model is used to calculate a risk score; the risk score is used to quantify the underwriting risk level. S2. Package the trained underwriting risk control model, preprocessing module, and FastAPI interface service script into a Docker image; In this embodiment of the invention, the trained underwriting risk control model and its dependent environment are packaged into a Docker image, which contains: Model weight file (e.g., .pkl or .h5 format); Preprocessing module (for feature normalization / encoding); FastAPI interface service script.

[0049] The preprocessing module includes: A classification feature encoder is used to convert textual medical data into numerical features that can be recognized by the model. Numerical normalizer processes continuous features according to a set formula; Missing value handler: fills in preset default values ​​for unprovided feature fields.

[0050] The FastAPI interface service script includes: The route definition module is used to declare the / risk_score and / decision interface paths; The request verification module verifies the integrity and format of the input data fields. The service startup module binds the interface to a specified port (such as 8000) within the container via UVicorn.

[0051] The Docker image building process can be achieved using the following Dockerfile: Dockerfile FROM python:3.8-slim WORKDIR / app # 1. Copy the model file, preprocessing module, and FastAPI script. COPY model.pkl preprocess.py api_server.py . / # 2. Install dependencies RUN pip install fastapi uvicorn scikit-learn # 3. Declare the startup command CMD ["uvicorn", "api_server:app", "--host", "0.0.0.0", "--port", "8000"] The packaging process is as follows: Save the trained underwriting risk control model as a serialized file (e.g., .pkl or .onnx format); the preprocessing module is implemented as a Python script (preprocess.py), which includes feature encoding, standardization, and missing value handling methods; the FastAPI interface service script (api_server.py) defines the / risk_score route, integrating model invocation and preprocessing logic.

[0052] Create a requirements.txt file to declare the dependent libraries and their versions; Build based on the official Python image, copy files and install dependencies; Execute the build command to generate the image: bash docker build -t risk-model:v1.0 . Start a temporary container to verify service availability: bash docker run -p 8000:8000 risk-model:v1.0.

[0053] S3. Deploy Docker images as multi-instance services (i.e., containers) through a Kubernetes cluster and configure them as follows: When each container starts, the FastAPI service automatically binds to a specified port within the container and listens on that port; The container port is mapped to a unified entry point for the cluster to receive external requests through the Service resource; Deploy this image as a scalable microservice instance (container) using Kubernetes, with each instance listening on a specified port; A RESTful risk control interface (FastAPI interface) is built based on FastAPI. The interface routing includes: / risk_score (Receives insurance applications and returns a risk score); / decision (Returns underwriting recommendations based on the score).

[0054] Starting the FastAPI interface service script includes: Load FastAPI application instances via the UVicorn server; Explicitly bind to 0.0.0.0 and a specified port (e.g., 8000) to allow external access; Declare the same port number in the Docker startup command.

[0055] The specific steps for deploying Docker images for multi-instance services via a Kubernetes cluster include: Create a Deployment resource file (deployment.yaml) to define the initial number of instances, image name, container listening port, etc. Create a Service resource file (service.yaml) and map the instance port to the unified entry point of the cluster; Execute the deployment command: bash kubectl apply -f deployment.yaml kubectl apply -f service.yaml Verify instance status: bash `kubectl get pods` # Check if the Pod is running. `kubectl get svc` # Get the external IP address of the Service Create a HorizontalPodAutoscaler resource (hpa.yaml) to enable automatic scaling.

[0056] S4. External requests are distributed to containers via Service load balancing. The port bound to the container receives the requests and routes them to the corresponding FastAPI interface, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score; the calculation result is returned through the original port path. It's important to note that each Docker container's FastAPI service needs to be bound to a port (e.g., 8000), which is the sole entry point for communication between the container and the outside world. Without port listening, the interface cannot receive requests. Kubernetes uses a Service to map the container port (e.g., 8000) to a unified cluster port (e.g., 80). External access only requires http: / / <cluster IP>:80, without needing to know the container's internal port.

[0057] The FastAPI service needs to declare a listening port when starting up, and can be accessed externally via http: / / <ip>Accessing the risk control interface via port 8000 / risk_score will route requests to the / risk_score path of FastAPI.

[0058] S5. Dynamically adjust the underwriting strategy based on real-time data feedback and returned risk scores: If the risk score exceeds the threshold, a decision to refuse coverage will be made. If the risk score is below the threshold, a low premium or standard premium recommendation will be given.

[0059] In some embodiments, the method further includes: Prometheus collects CPU utilization and FastAPI response time metrics for each container in real time. When any metric exceeds the corresponding preset threshold and remains so for a set time, Kubernetes HPA is triggered to automatically increase the number of containers. When all metrics are below half of the set threshold and remain so for a set time, Kubernetes HPA is triggered to automatically decrease the number of containers. in,

[0060]

[0061] In the formula, n is the total number of sampling requests.

[0062] Dynamically adjusting the number of instances includes: Define the CPU utilization threshold (70%) and the interface response time threshold (500ms). If any metric exceeds the limit for 30 seconds, HPA will gradually increase the number of instances from a minimum of 2 to a maximum of 10; when the metric falls below the threshold, it will reduce the number of instances by 1 every 5 minutes until it reaches the minimum. When Prometheus detects that the average response time of the interface exceeds 500ms, it will automatically increase the number of model instance replicas (up to a maximum of 10); it will send a / health probe request to the container every 5 seconds, and restart the container if it fails 3 times in a row; faulty containers will be automatically removed from the load balancer pool.

[0063] During adjustment,

[0064] For example, if the current CPU utilization is 85%, the target threshold is 70%, and the current number of containers is 2, the expected number of containers = ⌈85 / 70×2⌉ = 3. To avoid frequent fluctuations, the scaling down action is delayed by 5 minutes.

[0065] In some embodiments, step S1, which involves jointly training the underwriting risk control model using XGBoost and a neural network, includes: S11. Preprocess historical underwriting data, customer information, and risk characteristics; S12. Use the XGBoost model to train the preprocessed data. The XGBoost model automatically calculates the importance score of each feature during the training process. Obtain the importance score of each feature, which represents the degree of contribution of the feature to the model. S13. Sort all features according to their importance scores and select features with importance scores higher than a preset threshold as key features. S14. Construct a neural network model based on the selected key features; Here, we construct a three-layer fully connected neural network with the following structure: Input layer: XGBoost predicted probabilities + original features; Hidden layer: 128 neurons (ReLU activated); Output layer: 1 neuron (Sigmoid activation); Using the Adam optimizer with an initial learning rate of 0.001; Early stopping strategy (patience=10) prevents overfitting; S15. Input the output features of the XGBoost model together with the preprocessed original features into the neural network model for training, and adjust the model parameters using the Adam algorithm to minimize the loss function. loss function

[0066] loss function Measure the degree of difference between the model's predictions and the true labels; Indicates the first i The true label of each sample The model represents the first i The predicted probability of a sample. This represents the total number of samples.

[0067] S15. The prediction results of the XGBoost model and the neural network model are fused to obtain the final underwriting risk control model prediction result; when the prediction result meets the set requirements, the model training is completed.

[0068] The conditions for completing model training include: Before training the model, set thresholds for model performance metrics based on business needs, such as accuracy, recall, and F1 score. During training, the performance of the model is monitored using a validation set, and the performance metrics of the model on the validation set are calculated. The model is considered to have completed training when its performance metrics on the validation set reach or exceed a set threshold. If, after a certain number of training iterations, the model's performance metrics fail to reach the set threshold and no longer improve significantly, an early stopping mechanism is triggered to stop training in order to avoid overfitting and save computational resources.

[0069] It should be noted that the steps for calculating the importance score of each feature during XGBoost model training include: During the training of each tree, for each split point of each feature, the gain brought by the split is calculated; For each feature, the gains of that feature in all trees are summed to obtain the total gain of the feature; The importance score of a feature is obtained by dividing its total gain by the sum of the total gains of all features. The specific calculation formula is as follows:

[0070] In the formula, s represents the total number of trees in the model. m This represents the total number of features.

[0071] By calculating the gain from each feature split during the training of each tree, and then summing and normalizing the gains of that feature across all trees, an importance score for the feature is obtained. This method objectively evaluates the contribution of each feature to the model, providing an accurate basis for feature selection, and helps optimize the model structure, improve the model's generalization ability, and predictive accuracy.

[0072] The gain calculation method includes: for each split point of each feature, calculating the Gini impurity before and after splitting, and calculating the gain based on the Gini impurity before and after splitting;

[0073]

[0074] In the formula, Indicates the first The proportion of samples of each class in the dataset, where Q represents the total number of classes. For the split subset of data The number of samples, Data set before splitting The number of samples.

[0075] Gini impurity is a commonly used metric for measuring dataset purity. By calculating the change in Gini impurity before and after splitting, we can accurately assess the extent to which feature splitting improves model performance, thereby more reasonably calculating feature importance scores, further optimizing the model training process, and improving the model's prediction performance.

[0076] In some embodiments, step S4 specifically includes: Once an external request arrives at the Kubernetes Service, it is load-balanced and distributed to the containers. The container receives requests through the port bound to the FastAPI service, and FastAPI routes them to the corresponding interface processing function. The interface processing function parses the request to obtain the required data, triggers the preprocessing module in the container to parse the data, and calls the underwriting risk control model to calculate the risk score. The calculation results will be returned via the original port path.

[0077] The complete process after the interface receives a request includes: Request validation: Validate required fields (such as customer_id) and data format (such as medical_history being a JSON array); The preprocessing module is invoked to encode the raw data (e.g., convert disease names to ICD-10 codes); the numerical features are normalized (e.g., the sum insured is divided by 1 million).

[0078] The processed feature vector is input into the underwriting risk control model, and the model.predict() method is called to calculate the risk score; the model output is post-processed to convert it into a 0~1 standardized score.

[0079] Results returned: The response body includes a risk score, decision recommendations, and model version number (e.g., v1.2.3).

[0080] In some embodiments, the steps of the interface processing function parsing the request to obtain the required data, triggering the preprocessing module in the container to parse the data, and calling the underwriting risk control model to calculate the risk score include: The interface processing function extracts the customer's basic information, historical claims records, and health status characteristics from the request; it then triggers the preprocessing module within the container to preprocess the extracted features to make them conform to the format requirements of the model input. The interface receives request data in JSON format, with fields including: customer_id (unique customer identifier); medical_history (medical record code); insurance_type (insurance type).

[0081] The preprocessing module preprocesses features: it performs One-Hot encoding on categorical features (such as disease codes) and standardizes numerical features (such as age and insured amount) using Z-Score.

[0082] The extracted features are input into the trained XGBoost model to obtain the intermediate feature representation output by the XGBoost model; at the same time, the preprocessed original features are combined with the intermediate features output by XGBoost to form an enhanced feature set. The enhanced feature set is input into the trained neural network model, and the risk score for each customer is calculated through the forward propagation of the neural network model. The calculated risk scores are then standardized. Among them, risk score

[0083] In the formula, For the first Each feature weight, For the first There are 1 feature values, where m is the total number of features.

[0084] like The result is {"decision": "reject", "score": 0.85}. like It returns {"decision": "manual_review", "score": 0.45}; like The result is {"decision": "accept", "score": 0.2}.

[0085] For requests with the same customer_id, the result is read from the Redis cache first, with the cache key being risk_score:{customer_id} and an expiration time of 10 minutes; if the cache is not hit, the model is called to calculate the result, which is then written to Redis and returned.

[0086] like Figure 2 As shown, this embodiment of the invention also provides a real-time underwriting risk control system based on containerized deployment, including: The model training module, based on historical underwriting data, uses XGBoost and neural networks to jointly train the underwriting risk control model. The underwriting risk control model is used to calculate the risk score; the risk score is used to quantify the underwriting risk level. The containerization deployment module is used to encapsulate the trained underwriting and risk control model, preprocessing module, and FastAPI interface service script into a Docker image; deploy the Docker image as a multi-instance service, i.e., a container, through a Kubernetes cluster, and configure it so that when each container starts, the FastAPI service automatically binds to the specified port inside the container and listens on the port; and map the container port to a unified cluster entry point to receive external requests through the Service resource. The risk control interface engine distributes external requests to containers via Service load balancing. The containers receive requests on their bound ports and route them to the corresponding FastAPI interfaces, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score. The calculation results are then returned through the original port path. The dynamic decision-making module is used to dynamically adjust the underwriting strategy based on real-time data feedback and the returned risk score: when the risk score exceeds the threshold, a decision to refuse underwriting is returned; when the risk score is below the threshold, a suggestion of low premium or standard premium is returned.

[0087] In some embodiments, the containerization deployment module is also used to collect CPU utilization and FastAPI response time metrics of each container in real time through Prometheus; when any metric exceeds the corresponding preset threshold and continues for a set time, Kubernetes HPA is triggered to automatically increase the number of containers; when all metrics are below half of the set threshold and continue for a preset time, Kubernetes HPA is triggered to automatically decrease the number of containers. in,

[0088]

[0089] In the formula, n is the total number of sampling requests.

[0090] In some embodiments, the model training module includes a data preprocessing unit, a first-order training unit, a second-order training unit, and a training processing unit. The data preprocessing unit is used to preprocess historical underwriting data, customer information, and risk characteristics; The first-order training unit is used to train the preprocessed data using the XGBoost model. During the training process, the XGBoost model automatically calculates the importance score of each feature. The importance score of each feature is obtained, which represents the contribution of the feature to the model. All features are sorted according to their importance scores, and features with importance scores higher than a preset threshold are selected as key features. The second-order training unit is used to build a neural network model based on the selected key features. The output features of the XGBoost model and the preprocessed original features are input into the neural network model for training. The model parameters are adjusted by the Adam algorithm to minimize the loss function. The training processing unit is used to fuse the prediction results of the XGBoost model and the neural network model to obtain the final underwriting risk control model prediction result; when the prediction result meets the set requirements, the model training is completed.

[0091] In some embodiments, the first-order training unit is also used to calculate the gain brought by splitting at each split point of each feature during the training process of each tree; for each feature, the gains of that feature in all trees are accumulated to obtain the total gain of the feature; the total gain of the feature is divided by the sum of the total gains of all features to obtain the importance score of the feature, and the specific calculation formula is as follows:

[0092] In the formula, s represents the total number of trees in the model. m This represents the total number of features.

[0093] In some embodiments, the gain calculation method includes: for each split point of each feature, calculating the Gini impurity before and after splitting, and calculating the gain based on the Gini impurity before and after splitting.

[0094]

[0095] In the formula, Indicates the first The proportion of class samples in the dataset Q Indicates the total number of categories. For the split subset of data The number of samples, Data set before splitting The number of samples.

[0096] In some embodiments, the risk control interface engine is specifically used to load balance and distribute external requests to various containers after they arrive at the Kubernetes Service. The containers receive requests through the port bound to the FastAPI service, and FastAPI routes them to the corresponding interface processing function. The interface processing function parses the request to obtain the required data, triggers the preprocessing module in the container to parse the data, calls the underwriting risk control model to perform risk scoring calculation, and returns the calculation result through the original port path.

[0097] In some embodiments, the interface processing function extracts the customer's basic information, historical claims records, and health status characteristics from the request; it then triggers the preprocessing module within the container to preprocess the extracted features to make them conform to the format requirements of the model input. The risk control interface engine inputs the extracted features into the trained XGBoost model to obtain the intermediate feature representation output by the XGBoost model; at the same time, it combines the preprocessed original features with the intermediate features output by XGBoost to form an enhanced feature set; the enhanced feature set is input into the trained neural network model, and the risk score for each customer is calculated through the forward propagation of the neural network model; the calculated risk score is then standardized. Among them, risk score

[0098] In the formula, For the first Each feature weight, For the first There are 1 feature values, where m is the total number of features.

[0099] This invention also provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The communication bus can be used for information transmission between the electronic device and sensors. The processor can call logical instructions in memory to execute the following methods: S1. Based on historical underwriting data, an underwriting risk control model is jointly trained using XGBoost and a neural network. The underwriting risk control model is used to calculate a risk score, which is used to quantify the underwriting risk level. S2. The trained underwriting risk control model, preprocessing module, and FastAPI interface service script are packaged into a Docker image. S3. The Docker image is deployed as a multi-instance service (container) through a Kubernetes cluster, and configured such that: when each container starts, the FastAPI service automatically binds to a specified port within the container and listens to that port; the container port is mapped to a unified cluster entry point to receive external requests through Service resources. S4. External requests are distributed to containers via Service load balancing. The port bound to the container receives the request and routes it to the corresponding FastAPI interface, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score; the calculation result is returned through the original port path. S5. The underwriting strategy is dynamically adjusted based on real-time data feedback and the returned risk score: when the risk score exceeds a threshold, a rejection decision is returned; when the risk score is below the threshold, a low premium or standard premium suggestion is returned.

[0100] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] This invention provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the methods provided in the above-described method embodiments. For example, the methods include: S1, training an underwriting risk control model using XGBoost and a neural network based on historical underwriting data; the underwriting risk control model is used to calculate a risk score; the risk score is used to quantify the underwriting risk level; S2, encapsulating the trained underwriting risk control model, preprocessing module, and FastAPI interface service script into a Docker image; S3, deploying the Docker image as a multi-instance service (container) through a Kubernetes cluster, and configuring each container... Upon startup, the FastAPI service automatically binds to a specified port within the container and listens on that port. The container port is mapped to a unified cluster entry point via Service resources to receive external requests. S4: External requests are distributed to the container via Service load balancing. The port bound to the container receives the request and routes it to the corresponding FastAPI interface, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score. The calculation result is returned via the original port path. S5: The underwriting strategy is dynamically adjusted based on real-time data feedback and the returned risk score: if the risk score exceeds a threshold, a rejection decision is returned; if the risk score is below the threshold, a low premium or standard premium suggestion is returned.

[0102] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the present invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the present invention by those skilled in the art without departing from the spirit and essence of the invention, and such modifications or substitutions should all be within the scope of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should also be covered within the protection scope of the present invention.< / ip>

Claims

1. A real-time underwriting risk control method based on containerized deployment, characterized in that, include: S1. Based on historical underwriting data, an underwriting risk control model is jointly trained using XGBoost and neural networks. The underwriting risk control model is used to calculate a risk score; the risk score is used to quantify the underwriting risk level. S2. Package the trained underwriting risk control model, preprocessing module, and FastAPI interface service script into a Docker image; S3. Deploy Docker images as multi-instance services (i.e., containers) through a Kubernetes cluster and configure them as follows: When each container starts, the FastAPI service automatically binds to a specified port within the container and listens on that port; The container port is mapped to a unified entry point for the cluster to receive external requests through the Service resource; S4. External requests are distributed to containers via Service load balancing. The port bound to the container receives the requests and routes them to the corresponding FastAPI interface, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score; the calculation result is returned through the original port path. S5. Dynamically adjust the underwriting strategy based on real-time data feedback and returned risk scores: If the risk score exceeds the threshold, a decision to refuse coverage will be made. If the risk score is below the threshold, a low premium or standard premium recommendation will be given.

2. The real-time underwriting risk control method based on containerized deployment according to claim 1, characterized in that, The method also includes: Prometheus collects CPU utilization and FastAPI response time metrics for each container in real time. When any metric exceeds the corresponding preset threshold and remains so for a set time, Kubernetes HPA is triggered to automatically increase the number of containers. When all metrics are below half of the set threshold and remain so for a set time, Kubernetes HPA is triggered to automatically decrease the number of containers. in, In the formula, n is the total number of sampling requests.

3. The real-time underwriting risk control method based on containerized deployment according to claim 1, characterized in that, Step S1, which involves jointly training the underwriting risk control model using XGBoost and a neural network, includes the following steps: Preprocess historical underwriting data, customer information, and risk characteristics; The XGBoost model is used to train the preprocessed data. During the training process, the XGBoost model automatically calculates the importance score of each feature. The importance score of each feature represents the degree of contribution of the feature to the model. All features are sorted according to their importance scores, and features with importance scores higher than a preset threshold are selected as key features. A neural network model is constructed based on the selected key features; The output features of the XGBoost model are fed into the neural network model along with the preprocessed original features for training. The model parameters are then adjusted using the Adam algorithm to minimize the loss function. The prediction results of the XGBoost model and the neural network model are fused to obtain the final underwriting risk control model prediction result; when the prediction result meets the set requirements, the model training is completed.

4. The real-time underwriting risk control method based on containerized deployment according to claim 3, characterized in that, The steps for calculating the importance score of each feature during XGBoost model training include: During the training of each tree, for each split point of each feature, the gain brought by the split is calculated; For each feature, the gains of that feature in all trees are summed to obtain the total gain of the feature; The importance score of a feature is obtained by dividing its total gain by the sum of the total gains of all features. The specific calculation formula is as follows: In the formula, s represents the total number of trees in the model. m Indicates the total number of features.

5. The real-time underwriting risk control method based on containerized deployment according to claim 4, characterized in that, The gain calculation method includes: for each split point of each feature, calculating the Gini impurity before and after splitting, and calculating the gain based on the Gini impurity before and after splitting; In the formula, Indicates the first The proportion of samples of each class in the dataset, where Q represents the total number of classes. For the split subset of data The number of samples, Data set before splitting The number of samples.

6. The real-time underwriting risk control method based on containerized deployment according to claim 1, characterized in that, Step S4 specifically includes: Once an external request arrives at the Kubernetes Service, it is load-balanced and distributed to the containers. The container receives requests through the port bound to the FastAPI service, and FastAPI routes them to the corresponding interface processing function. The interface processing function parses the request to obtain the required data, triggers the preprocessing module in the container to parse the data, and calls the underwriting risk control model to calculate the risk score. The calculation results will be returned via the original port path.

7. The real-time underwriting risk control method based on containerized deployment according to claim 6, characterized in that, The steps of the interface processing function parsing the request to obtain the required data, triggering the preprocessing module in the container to parse the data, and calling the underwriting risk control model to calculate the risk score include: The interface processing function extracts the customer's basic information, historical claims records, and health status characteristics from the request; it then triggers the preprocessing module within the container to preprocess the extracted features to make them conform to the format requirements of the model input. The extracted features are input into the trained XGBoost model to obtain the intermediate feature representation output by the XGBoost model; at the same time, the preprocessed original features are combined with the intermediate features output by XGBoost to form an enhanced feature set. The enhanced feature set is input into the trained neural network model, and the risk score for each customer is calculated through the forward propagation of the neural network model. The calculated risk scores are then standardized. Among them, risk score In the formula, For the first Each feature weight, For the first There are 1 feature values, where m is the total number of features.

8. A real-time underwriting risk control system based on containerized deployment, characterized in that, include: The model training module, based on historical underwriting data, uses XGBoost and neural networks to jointly train the underwriting risk control model. The underwriting risk control model is used to calculate the risk score; the risk score is used to quantify the underwriting risk level. The containerization deployment module is used to encapsulate the trained underwriting and risk control model, preprocessing module, and FastAPI interface service script into a Docker image; deploy the Docker image as a multi-instance service, i.e., a container, through a Kubernetes cluster, and configure it so that when each container starts, the FastAPI service automatically binds to the specified port inside the container and listens on the port; and map the container port to a unified cluster entry point to receive external requests through the Service resource. The risk control interface engine distributes external requests to containers via Service load balancing. The containers receive requests through their bound ports and route them to the corresponding FastAPI interfaces, triggering the preprocessing module to parse the data and call the underwriting risk control model to calculate the risk score. The calculation results will be returned via the original port path; The dynamic decision-making module is used to dynamically adjust the underwriting strategy based on real-time data feedback and returned risk scores: when the risk score exceeds the threshold, a decision to refuse underwriting is returned; If the risk score is below the threshold, a low premium or standard premium recommendation will be given.

9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, the computer program instructions being executed by the at least one processor to enable the at least one processor to execute the real-time underwriting and risk control method based on containerized deployment as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the real-time underwriting and risk control method based on containerized deployment as described in any one of claims 1 to 7.