Insurance product pricing method and device based on data mining, equipment and medium

By constructing an insurance product pre-quotation model using federated learning and differential privacy technology, the problems of insufficient model generalization ability and credibility in insurance pre-quotation methods are solved, enabling fast and accurate pre-quotations and improving the market responsiveness and customer trust of insurance companies.

CN122390877APending Publication Date: 2026-07-14CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA PING AN PROPERTY INSURANCE CO LTD
Filing Date
2026-03-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing insurance pre-quotation methods suffer from problems such as insufficient model generalization ability, insufficient reliability of quotation results, and low model training efficiency.

Method used

By acquiring local training data from multiple participants, encrypting it using a privacy-preserving protocol, and then performing distributed modeling within a federated learning framework, a global model is constructed. The model parameters are then processed using differential privacy technology, and the global model parameters are updated using a secure aggregation algorithm, thereby enabling differentiated insurance product pre-pricing.

Benefits of technology

It improves the adaptability and prediction accuracy of the pre-quotation model to different customer groups, enhances the credibility of the quotation results and the practicality of the system, and solves the problems of slow response, high information threshold and easy missed business opportunities in the traditional quotation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122390877A_ABST
    Figure CN122390877A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of artificial intelligence, and discloses an insurance product price forecasting method, device, equipment and medium based on data mining, which is applied to the fields of financial technology and medical health. The method comprises the following steps: obtaining a historical insurance policy dataset; performing clustering analysis on the historical insurance policy dataset by using a clustering algorithm; based on the result of the clustering analysis, screening a feature set related to premium data from dimensional feature data and claim data, which satisfies a preset condition; constructing a price forecasting model according to the feature set; in response to a price forecasting request for a target object, obtaining enterprise attribute information of the target object and its associated historical business data; inputting the enterprise attribute information and the historical business data into the price forecasting model, and outputting reference pricing information of the target object. The application solves the problems existing in the prior art of insurance product price forecasting methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology and is applied in the fields of financial technology and healthcare. In particular, it relates to a method, apparatus, equipment and medium for pre-quoting insurance products based on data mining. Background Technology

[0002] As the digital transformation of the insurance industry deepens, leveraging data-driven methods to achieve rapid and accurate pre-quotations for insurance products has become crucial for insurance companies to improve market responsiveness and seize business opportunities. However, insurance operations face significant data and efficiency challenges in the actual quoting process. On the one hand, the limited number of underwriters makes it difficult to quickly respond to the large volume of immediate quotation requests from front-end agents; on the other hand, potential clients are often unwilling or unable to provide detailed company information in the early stages of a business opportunity, causing traditional actuarial or rule-based models to fail to activate effectively due to insufficient input features. Furthermore, existing quotation models typically rely on a large number of complex inquiry factors, resulting in excessively long information collection times for agents, making them highly vulnerable to being outmaneuvered by faster competitors and losing business opportunities.

[0003] Artificial intelligence technology offers a potential solution to the aforementioned problems. Theoretically, by mining patterns in historical policy data and building data-driven predictive models, reference quotes can be output with limited input information, enabling rapid self-service quoting. Existing solutions typically rely on historical data to directly train regression or classification models to predict premiums.

[0004] However, in practical applications, existing data-driven pricing methods have significant shortcomings, particularly in terms of processing efficiency, data generalization, and model usability. First, in the data preprocessing and feature engineering stages, traditional methods flatten historical policy data, failing to effectively perform differentiated clustering and labeling based on industry, scale, risk, and other dimensions. This makes it difficult for the model to capture the differentiated risk-premium relationships among different customer groups, resulting in insufficient generalization ability. Second, in the model training stage, differentiated validation strategies (such as cross-validation and hold-out methods) are not adopted for customer groups with different data volumes (e.g., niche industries versus mass industries). This makes it impossible to guarantee model stability in data-scarce scenarios or to efficiently utilize data in data-rich scenarios. Furthermore, in the pricing generation and credibility building stages, existing methods typically only output an isolated value or range, lacking the tracing and display of similar historical underwriting cases, reducing the trust of agents and customers in the pre-quotation results. Summary of the Invention

[0005] This invention provides a data mining-based method, apparatus, equipment, and medium for insurance product pre-quotation, to address the problems of insufficient model generalization ability, insufficient reliability of quotation results, and low model training efficiency in existing insurance pre-quotation methods.

[0006] In a first aspect, the present invention provides a method for pre-quoting insurance products based on data mining, comprising: The system acquires local training data from multiple participants, encrypts the original data of each local training data using a privacy protection protocol, and generates a corresponding distributed model by performing distributed modeling on each encrypted local training data using a federated learning framework. Configure the training parameters for the federated learning task and construct a global model based on multiple distributed models. The parameters of the global model are processed based on differential privacy technology to generate privacy-protected global model parameters. The privacy-protected global model parameters are aggregated and calculated using a secure aggregation algorithm, and the global model parameters are updated based on the aggregation results.

[0007] Secondly, the present invention provides an insurance product pre-quotation device based on data mining, comprising: The data acquisition module is used to acquire local training data from multiple participants, encrypt the original data of each local training data through a privacy protection protocol, and generate a corresponding distributed model by performing distributed modeling on each encrypted local training data through a federated learning framework. The model building module configures the training parameters for the federated learning task and builds a global model based on multiple distributed models. The parameter generation module processes the parameters of the global model based on differential privacy technology to generate privacy-protected global model parameters. The parameter update module performs aggregate calculations on the privacy-protected global model parameters using a secure aggregation algorithm, and updates the global model parameters based on the aggregation results.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described data mining-based insurance product pre-quotation method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described data mining-based insurance product pre-quotation method.

[0010] The aforementioned solution, implemented through data mining-based insurance product pre-quotation methods, devices, equipment, and media, utilizes data-driven historical policy clustering analysis and differentiated modeling to group policies based on multi-dimensional characteristics such as industry and scale, and extracts risk-premium relationships. This effectively overcomes the insufficient generalization ability caused by the flattening of data in traditional methods. Through differentiated verification strategies and feature engineering for high and low data volume groups, the solution ensures model stability in data-scarce scenarios while fully exploring potential patterns in rich data scenarios, significantly improving the adaptability and prediction accuracy of the pre-quotation model for different customer groups. This solution enables sales agents to quickly and independently obtain relatively accurate reference quotes under conditions of limited customer information and strained underwriting resources, technically resolving the core contradictions of slow response, high information barriers, and missed business opportunities inherent in traditional quotation processes.

[0011] Furthermore, the core reason for the insufficient credibility and practicality of pre-quotation results in related technologies is the isolated output information and lack of business context. This solution, however, while generating a reference quotation range, simultaneously queries external company information and internal underwriting and claims databases, automatically matching and presenting similar historical underwriting cases that have undergone anonymization. This places the pre-quotation results within a real business reference system, providing sales personnel with an intuitive and credible basis for communication. This not only improves the acceptability of quotations but also enhances the professional image and responsiveness of the sales team, fundamentally improving the practicality and persuasiveness of the pre-quotation system in real sales scenarios.

[0012] In summary, this solution can address the problems of insufficient model generalization ability, insufficient reliability of pricing results, and low model training efficiency in existing insurance pre-quotation methods. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart illustrating a data mining-based insurance product pre-quotation method according to an embodiment of the present invention.

[0015] Figure 2 yes Figure 1 A flowchart of step S120.

[0016] Figure 3 yes Figure 1 A flowchart of step S130.

[0017] Figure 4 yes Figure 1 A flowchart of step S140.

[0018] Figure 5 yes Figure 1 A flowchart of step S150.

[0019] Figure 6 yes Figure 1 Another flowchart of step S140.

[0020] Figure 7 This is a flowchart illustrating a data mining-based insurance product pre-quotation method according to an embodiment of the present invention.

[0021] Figure 8 yes Figure 7 A flowchart of step S170.

[0022] Figure 9 This is a schematic diagram of a data mining-based insurance product pre-quotation device according to an embodiment of the present invention.

[0023] Figure 10 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention.

[0024] Figure 11 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] Please see Figure 1 As shown in the flowchart, this embodiment of the invention provides a data mining-based method for pre-quoting insurance products, which includes the following steps.

[0027] Step S110: Obtain historical policy dataset, which includes policy dimensional feature data, premium data, and claims data.

[0028] Specifically, historical policy datasets can originate from the insurance company's core business systems, underwriting databases, and claims databases. This dataset aims to provide comprehensive and structured historical experience for subsequent model training. In a healthcare implementation example: taking a health insurance company as an example, when developing a pre-quotation model for a "group high-end medical insurance" product, the historical policy dataset specifically includes: dimensional feature data: the insured company's industry category (e.g., information technology, biopharmaceuticals, high-end manufacturing), total number of employees, average age distribution of employees, historical annual physical examination abnormality statistics (e.g., detection rate of hypertension and hyperlipidemia), selected insurance product plan level (e.g., basic version, premium version), coverage area, etc.; premium data: the total annual premium corresponding to each historical policy; claims data: claims records for each historical policy during the coverage period, including claim type (e.g., outpatient, inpatient, special drugs), total claim amount, number of claims, and types of common diseases, etc. For example, a historical insurance policy from a biopharmaceutical company has the following dimensional characteristics: "Industry: Biopharmaceutical; Number of Employees: 1500; Average Age: 35; Hypertension Detection Rate: 8%; Product Plan: Premium Edition." The corresponding annual premium is 3 million yuan, and there were 12 hospitalization claims that year, with a total payout of 800,000 yuan. All of these structured data constitute a training sample. In an example from a fintech implementation: taking a property insurance company as an example, when developing a pre-quotation model for "comprehensive property insurance for micro and small enterprises," the historical policy dataset specifically includes: dimensional feature data: the enterprise's industry (e.g., catering and retail, software services, logistics and transportation), registered capital, nature of business premises (e.g., street-front shops, office buildings, industrial parks), building structure level, main inventory types, existing security protection measures (e.g., fire protection systems, security monitoring), etc.; premium data: the annual premium corresponding to each historical policy; claims data: all claims records that occurred during the coverage period of each historical policy, including the cause of the accident (e.g., fire, theft, water damage and pipe bursts), the amount of loss, the frequency of claims, etc. For example, a historical insurance policy from a restaurant chain has the following dimensional characteristics: "Industry: Catering; Registered Capital: 1 million yuan; Location: Street-front shop; Building Structure: Brick-concrete; Inventory: Food raw materials; Security: Basic fire protection." The corresponding annual premium is 12,000 yuan. Within the past three years, there was one claim for property damage caused by a fire in the kitchen exhaust duct, with a payout of 50,000 yuan. This record is extracted and included in the historical insurance policy dataset.

[0029] Step S120: Use a clustering algorithm to perform cluster analysis on the historical policy dataset to obtain multiple policy groups with different risk characteristics.

[0030] It should be noted that the clustering analysis in step S120 is a data structuring and insight discovery method based on unsupervised machine learning. Its core objective is to automatically classify massive and complex raw insurance policy data according to their inherent risk characteristics. The clustering process is closely related to the risk drivers in different insurance business areas. In the healthcare field, the risk factors affecting premiums are complex and diverse, including the disease spectrum, age structure, and medical habits of the insured population. The clustering algorithm integrates these multi-dimensional characteristics to group policyholders with similar health risk status (such as young tech company employees and older manufacturing workers). In the property insurance (fintech) field, risk is strongly correlated with the physical attributes, industry attributes, and historical claims records of the insured. The algorithm analyzes these characteristics to group companies with similar risk characteristics (such as multiple restaurants or multiple software companies) into the same group. This grouping is the basis for subsequent differentiated modeling and accurate pricing.

[0031] In some embodiments of the present invention, such as Figure 2 As shown, step S120 includes the following steps: Step S121: Extract multiple preset dimensional features from the historical policy dataset; Step S122: Use an unsupervised clustering algorithm to perform operations on the extracted dimensional features to aggregate insurance policies with similar risk characteristics into the same group; Step S123: Count the number of policies contained in each group, and divide the group into high data volume group and low data volume group based on the number.

[0032] Specifically, in step S121, the method for extracting preset dimensional features is based on defining and screening key risk factors using domain knowledge. These features are input variables for subsequent clustering and modeling. In the field of healthcare insurance, the extracted dimensional features typically include: the industry category of the insured company (e.g., finance, IT, manufacturing), the average age of the company's employees, the gender ratio, the historical annual abnormality rate of physical examinations (e.g., detection rate of hypertension, hyperglycemia), the selected coverage area (domestic / global), and the product plan type (basic / premium). For example, from a group premium medical insurance policy underwritten for an internet company, the extracted feature vector might be: [Industry = Internet Technology, Average Age = 29 years old, Gender Ratio = 7:3, Hyperlipidemia Detection Rate = 5%, Coverage Area = Global, Product Plan = Premium]. In the field of property insurance (fintech), the extracted features include: the sub-industry of the company (e.g., catering, retail wholesale, software development), registered capital, the area and nature of the business premises (street-front / industrial park), the main property type (equipment / inventory), and the number of claims in the past three years. For example, the feature vector extracted from the property insurance policy of a small restaurant might be: [Industry = Catering, Registered Capital = 500,000, Location Type = Street-front Shop, Area = 200 square meters, Property Type = Kitchen Equipment Inventory, Number of Claims in the Past Three Years = 1].

[0033] Specifically, in step S122, the unsupervised clustering algorithm used for computation can be K-Means, DBSCAN, or hierarchical clustering, etc. Taking the K-Means algorithm as an example, its goal is to assign N policy samples to K clusters, minimizing the sum of squared distances from each sample to its cluster center. The algorithm first randomly initializes K cluster centers, and then iteratively executes the following two steps until convergence: 1) Assign each policy sample to the nearest cluster center; 2) Recalculate the mean of all samples in each cluster as the new cluster center. In the field of healthcare insurance, after clustering, a cluster characterized by "advanced age, manufacturing industry, and high incidence of chronic diseases" may be generated. This cluster contains a large number of employee medical insurance policies from similar heavy industrial enterprises, whose historical payout rates are generally high. At the same time, another cluster characterized by "young people, IT industry, and excellent health checkup indicators" may be generated, which contains policies from multiple internet companies, whose historical payout rates are low. In the property insurance field, clustering operations may aggregate numerous policies for "street-front shops, catering industry, and areas less than 300 square meters" into a high-risk cluster (high risk of fire and theft), while aggregating policies for "office buildings, software development industry, and no physical inventory" into a low-risk cluster.

[0034] Specifically, in step S123, the method of dividing into high and low data volume groups is based on a preset policy quantity threshold. The number of historical policies contained in each cluster is counted. A threshold M is set (e.g., M=1000). If the number of policies in a cluster is greater than or equal to M, it is marked as a "high data volume group"; if it is less than M, it is marked as a "low data volume group". This division is crucial for subsequent model training strategies. In the field of healthcare insurance, a cluster containing thousands of group medical insurance policies from "large enterprises with a wide age distribution of employees" is classified as a high data volume group, with sufficient data suitable for training complex models; while a cluster containing only a few hundred medical insurance policies for employees of "rare disease specialty hospitals" is classified as a low data volume group, requiring a special training strategy to prevent overfitting. In the field of property insurance, a cluster with tens of thousands of property insurance policies for "national chain retail supermarkets" belongs to the high data volume group; while a cluster with only a hundred or so special property insurance policies for "new energy battery manufacturing plants" belongs to the low data volume group. Due to the scarcity of samples, the risk patterns of the latter are more difficult to learn.

[0035] Understandably, unsupervised clustering analysis and grouping mechanisms based on historical policy data, compared to traditional methods that process all data using a single model, can more precisely identify and differentiate customer groups with different risk attributes, laying a data foundation for building differentiated and precise pricing models. In the healthcare insurance field, this mechanism is crucial because accurately identifying different risk clusters, such as "elderly groups with high rates of chronic diseases" and "young people with excellent health," helps insurance companies design more risk-targeted products and rates for different customer groups, providing data-driven insights for underwriting and precision marketing. This clustering-based customer segmentation framework has a powerful ability to identify complex and diverse health risk factors. Whether based on physiological indicators from physical examination data or socioeconomic factors based on industry and region, the algorithm can automatically discover inherent group patterns. Unlike traditional coarse segmentation based on simple rules (such as only by age or industry), this mechanism significantly improves the scientific rigor and granularity of risk grouping through multi-dimensional feature-based clustering. After grouping policies with distinct risk characteristics, it provides clear guidance for subsequent differentiated modeling strategies for different data volumes. For example, when quoting prices, for high-data-volume groups (such as medical insurance for ordinary corporate employees), the rich data can be used to train a high-precision model; for low-data-volume groups (such as accident insurance for special occupational groups), few-shot learning techniques can be used to improve the model's generalization ability on sparse data, thereby improving the overall coverage and accuracy of the pre-quotation system across different markets, effectively meeting the needs of insurance companies for refined risk management and rapid market response for massive, heterogeneous customer groups. In the property insurance (fintech) field, this clustering analysis and grouping mechanism can accurately distinguish between different target clusters such as "high-risk street-front shops" and "low-risk office enterprises," helping insurance companies implement more scientific risk classification and differentiated pricing. This data-driven grouping technology demonstrates a strong ability to integrate and analyze the diverse attributes (physical, environmental, and operational) of property risks. Unlike traditional methods that mainly rely on industry categories and insured amounts for pricing, this mechanism, through multi-dimensional clustering, can discover the risk differences brought about by different business models within the same industry, significantly improving the depth of risk identification and the fairness of pricing. After grouping and segmenting, a structured input is provided for building a more robust pricing model. For example, in the underwriting process, more stringent underwriting procedures or additional conditions can be automatically triggered for identified high-risk clusters; while for low-risk clusters, rapid and automatic underwriting can be achieved. This improves the risk screening efficiency and cross-channel business collaboration capabilities of insurance companies, effectively meeting the urgent needs of modern insurance businesses for automated and precise risk pricing.

[0036] Step S130: Process the parameters of the global model based on differential privacy technology to generate privacy-protected global model parameters.

[0037] It should be noted that the feature selection and derivation in step S130 is a feature engineering process driven by statistical analysis and domain knowledge. Its core objective is to automatically identify the key risk factors with the strongest explanatory and predictive power for premiums from the original numerous dimensions within each identified policy group with distinct risk characteristics, and to uncover potential deep-seated risk interaction patterns through creative feature combinations. This process is deeply coupled with the pricing logic of different insurance sectors. In the healthcare insurance sector, this step aims to extract core driving features strongly correlated with medical expenses (such as the prevalence of specific chronic diseases) from complex data such as the health records, demographics, and protection plans of the insured population. In the property insurance (fintech) sector, it aims to locate essential features highly correlated with the probability and severity of property loss (such as the industry's inherent physical risk coefficient) from corporate attributes, asset status, and historical claims records.

[0038] In some embodiments of the present invention, such as Figure 3 As shown, step S130 includes the following steps: Step S131: For each policy group, calculate the correlation coefficient between each feature and the premium data within it; Step S132: Select features with correlation coefficients greater than the first threshold to form an initial feature subset; Step S133: Perform a combination operation on at least two features in the initial feature subset to generate at least one derived feature, and add the derived feature to the feature set.

[0039] Specifically, in step S131, the correlation coefficient can be calculated using methods such as Pearson correlation coefficient and Spearman rank correlation coefficient to quantify the strength of the linear or monotonic relationship between features and premiums. In the field of health insurance, for the high-risk group of "middle-aged and elderly manufacturing employees," the system will calculate the correlation coefficient between each feature (such as "average employee age," "historical hypertension detection rate," and "number of coverage items") of all historical policies within this group and the corresponding "annual average premium per person." The calculation may show that "historical hypertension detection rate" is strongly positively correlated with premiums (r=0.75), while the correlation of "city level where the company is registered" is very weak (r=0.08). In the field of property insurance (fintech), for the group of "warehousing and logistics companies," the calculation may find that "average value density of stored goods" is strongly positively correlated with premiums (r=0.68), "fire protection facility compliance level" is strongly negatively correlated with premiums (r=-0.61), while the correlation of "years of company establishment" is not significant (r=0.15).

[0040] Specifically, in step S132, the screening process is based on a preset threshold for the absolute value of the correlation coefficient (e.g., the first threshold is set to 0.3). For each group, only features with an absolute value of the correlation coefficient exceeding this threshold are retained, forming an initial feature subset for that group, thereby filtering out noise or irrelevant variables. In the field of healthcare insurance, for the low-risk group "young internet company employees," features such as "whether it includes maternity insurance" (r=0.55) and "annual physical examination package level" (r=0.45) may be selected, while "employee gender ratio" is excluded due to low correlation. In the field of property insurance, for the group "software development companies in office buildings," features such as "core code and data asset valuation" (r=0.72) and "whether it has purchased additional clauses for trade secret leakage insurance" (r=0.5) may be selected, while "office rental area" is excluded.

[0041] Specifically, in step S133, the derived features are generated by performing mathematical operations (such as arithmetic operations, ratios, products, etc.) on the elements of the initial feature subset to reveal the potential nonlinear impact of the interactions between features on premiums. In the field of health insurance, for the two selected features, "average employee age" and "prediabetes detection rate," a derived feature, "age-corrected diabetes risk index," can be generated, for example, defined as (average age / 45). Prediabetes detection rate. This derived feature may be a better predictor of future diabetes-related medical expenses than a single feature. In the property insurance sector, for the selected features "industry fire risk coefficient (external data)" and "company-owned fire brigade response time," a derived feature "theoretical maximum possible loss adjustment factor" can be generated, for example, defined as industry fire risk coefficient / log(response time + 1), to comprehensively measure disaster mitigation capabilities under specific risk exposures.

[0042] Understandably, this dynamic feature selection and intelligent derivation mechanism based on grouped statistical correlation analysis, compared to traditional methods that use fixed, empirical feature sets or apply uniform features to all data, can tailor the most relevant and predictive feature set for each risk segment. In the healthcare insurance field, this mechanism can automatically focus on health cost drivers specific to different population groups (such as chronic diseases in the elderly and specific protection needs of younger groups), and capture the risk aggregation effect through feature combinations, thus providing a clean and powerful input for building a high-precision differentiated pricing model, significantly improving the model's interpretability and pricing fairness. In the property insurance (fintech) field, this mechanism can accurately identify key risk variables in different industries and asset types, quantify the effectiveness of risk mitigation measures, and create composite risk indicators, enabling the model to more nuancedly assess complex and personalized risk situations.

[0043] Step S140: Construct a pre-quotation model based on the feature set.

[0044] It should be noted that the model building in step S140 is a data-driven, differentiated modeling process that fully considers the differences in data availability. Its core objective is to design and train the most suitable predictive sub-models for the different policy groups identified in step S120, especially those with vastly different data abundance, ultimately integrating them into a unified and intelligent pre-quotation system. This process demonstrates the practicality and targeted nature of model training. In the healthcare insurance field, this means employing different training strategies for "data-rich common disease groups" (such as employees of large enterprises) and "data-scarce special disease groups" (such as rare disease patient organizations); in the property insurance (fintech) field, it means implementing differentiated modeling schemes for "sample-rich common asset types" (such as ordinary shops) and "sample-scarce special asset types" (such as specialized laboratories).

[0045] In some embodiments of the present invention, such as Figure 4 As shown, step S140 includes the following steps: Step S141: For data that has been divided into high-volume single groups, the hold-out method is used to divide the training set and the validation set, and the first prediction sub-model is trained. Step S142: For data that has been divided into low-volume single groups, cross-validation is used to train the second prediction sub-model. Step S143: Integrate the first prediction sub-model and the second prediction sub-model to form the pre-quotation model.

[0046] Specifically, in step S141, for high-data-volume groups, a hold-out method is used for model training and validation. This method randomly divides the historical policy data within a group into two mutually exclusive parts; for example, 70% is used as the training set to learn model parameters, and 30% is used as an independent validation set to evaluate model performance and prevent overfitting. Due to the sufficient data volume, it can support training models with high complexity and strong expressive power. In the field of healthcare insurance, for the high-data-volume group of "Group Medical Insurance for Middle-aged White-collar Workers" containing tens of thousands of policies, a gradient boosting decision tree model (such as LightGBM) can be trained using the hold-out method (70 / 30 division) as the first predictive sub-model. This model can capture the complex nonlinear relationship between multiple features such as age, chronic disease history, and coverage scope and premiums. In the field of property insurance (fintech), for the group of "Vehicle Insurance for Nationwide Logistics Companies" with massive amounts of data, a deep neural network model is also trained using the hold-out method as the first predictive sub-model to learn deep patterns of multi-dimensional features such as vehicle type, driving route, and driver driving records, as well as claims costs.

[0047] Specifically, in step S142, for low-data-volume groups, cross-validation (such as K-fold cross-validation) is used to train the model. Due to data scarcity, to maximize the use of limited samples and obtain more robust performance estimates, all data is uniformly divided into K parts (e.g., K=5). One part is used as the validation set, and the remaining K-1 parts are used as the training set. Training and validation are repeated K times, and the average performance of the K iterations is taken as the model evaluation index, which is used to determine the final model parameters. For such groups, models with relatively simple structures and less prone to overfitting are usually chosen. In the field of health insurance, for the low-data-volume group of "Professional Athletes Group Accident and Health Insurance" with only a few hundred policies, 5-fold cross-validation is used to train a linear regression model with L1 / L2 regularization as the second predictive sub-model. Regularization constraints prevent overfitting on a small amount of data. In the field of property insurance, for the group of "Art Gallery Property All Risks Insurance" with only a small number of samples, 10-fold cross-validation is used to train a decision tree model (the complexity can be controlled by setting parameters such as maximum depth) as the second predictive sub-model.

[0048] Specifically, in step S143, the integration of various sub-models into a unified pre-quotation model involves constructing a model routing and scheduling framework. The core of this framework is a model scheduler based on group identifiers. When a new pre-quotation request arrives, the system first maps it to a specific policy group based on the input enterprise / customer characteristics, using the clustering model or rules trained in step S120. Then, the model scheduler automatically calls and executes the corresponding first or second prediction sub-model to predict the premium, based on the group's type label of "high data volume" or "low data volume". All sub-models are logically integrated into a complete "pre-quotation model," providing a unified prediction interface. In the healthcare insurance field, when inquiring about a quote for a large manufacturing company, the system categorizes it into the "high data volume - manufacturing employees" group and calls the corresponding LightGBM sub-model; when inquiring about a quote for an e-sports club, it categorizes it into the "low data volume - special occupational groups" group and calls the corresponding regularized linear regression sub-model. In the field of property insurance, a deep neural network sub-model is used to inquire about prices for a chain supermarket, while a decision tree sub-model is used to inquire about prices for an ancient building.

[0049] Understandably, this mechanism of differentiated modeling and intelligent integration based on data abundance, compared to traditional methods that use a single training paradigm for all data (such as holding out or cross-validation for all), can more scientifically and rationally address the data imbalance problem commonly found in insurance business. In the healthcare insurance sector, this mechanism ensures that high-precision, high-performance complex predictive models can be built by fully leveraging big data advantages for common mass customer groups, while stable and reliable predictive results can be obtained for niche and special customer groups through small-sample learning techniques. This expands market coverage while ensuring the overall robustness of the pricing model. In the property insurance (fintech) sector, this mechanism enables insurance companies to achieve accurate and automated risk pricing for mainstream target markets, while also providing feasible intelligent pricing support for long-tail, emerging, or special risk areas, enhancing the market adaptability and competitiveness of products. This flexible model architecture, combined with the aforementioned feature engineering and closed-loop optimization, constitutes the core of a continuously evolving, broadly covered, and accurately reliable intelligent pricing system, effectively solving the pain points of traditional methods where models fail in data-scarce scenarios or fail to fully exploit model potential in data-rich scenarios.

[0050] Step S150: In response to the pre-quotation request for the target object, obtain the enterprise attribute information of the target object and its associated historical business data.

[0051] It should be noted that step S150 is the real-time data preparation stage of the pre-quotation process. Its core lies in the rapid and automated aggregation of internal and external multi-source data to build a comprehensive information view for model inference for the target customer. This step addresses the pain points of inefficiency and incomplete information in traditional quoting processes where sales staff manually collect information. In the healthcare insurance field, this means not only obtaining the basic business registration information of the insured company but also querying the company's or similar companies' past group insurance policy underwriting and claims records to assess its overall health risk profile. In the property insurance (fintech) field, it requires integrating the target company's publicly disclosed business attributes with the historical risk performance of similar internal assets to comprehensively assess the target company's property risk exposure.

[0052] In some embodiments of the present invention, such as Figure 5 As shown, step S150 includes the following steps: Step S151: Obtain the publicly available enterprise information of the target object by calling an external enterprise information query interface; Step S152: Based on at least one key attribute in the publicly available corporate information, retrieve historical underwriting records with similar attributes from the internal underwriting database; Step S153: Retrieve relevant historical claims records from the internal claims database based on at least one key attribute in the publicly available corporate information.

[0053] Specifically, in step S151, the external enterprise information query interface is invoked through an application programming interface (API) to interact in real time with a third-party business information platform (such as "Tianyan Check," "Qichacha," or the official enterprise credit information disclosure system). By inputting the target enterprise's name or unified social credit code, structured or parsed public information can be obtained. In the field of healthcare insurance, obtaining public information about a company named "XX Biotechnology Co., Ltd." might include: its industry ("pharmaceutical manufacturing"), registered capital (RMB 50 million), number of insured persons (approximately 300), and business scope (including "biotechnology research and development"). This information forms the basis for assessing its industry risk and employee size. In the field of property insurance (fintech), obtaining information about a company named "YY Catering Management Co., Ltd." might include: its industry ("catering"), registered capital (RMB 1 million), registered address (a street in a certain district of a certain city), and number of branches (5). This information is crucial for assessing its business scale, geographical distribution, and the starting point of property risk.

[0054] Specifically, in step S152, the method for retrieving historical underwriting records from the internal underwriting database is based on fuzzy or exact matching of one or more key attributes. These key attributes are typically extracted from publicly available information obtained in step S151, such as "industry classification code," "company size (based on the number of insured persons)," and "registered region." Based on these attributes, the system searches the historical policy database for underwriting records of other companies with the same or highly similar attributes. In the field of healthcare insurance, for the aforementioned "XX Biotechnology Company," the system uses its industry (pharmaceutical manufacturing) and employee size (around 300 people) as key conditions to retrieve historical group health insurance policies of all companies in the same industry and of similar size, obtaining their coverage plans, historical premiums, and other data. In the field of property insurance, for "YY Catering Company," using its industry (catering) and business model (chain) as key conditions, the system retrieves historical property insurance policies of all similar catering companies, obtaining their insured amounts, rates, special terms, and other information.

[0055] Specifically, in step S153, the method of retrieving historical claims records from the internal claims database is similar to that in step S152, but the focus is on the actual occurrence of the risk. The system queries the historical claims database based on the same key attributes (such as company ID, industry, region, etc.) to obtain historical loss data for relevant companies. In the field of health insurance, for the same target company or a group of companies of similar size in the same industry, the system retrieves its historical medical claims records and analyzes statistical information such as the number of claims, high-frequency disease types, and average claim amount per person to quantify its health risk. In the field of property insurance, for the target catering company or similar companies, the system retrieves its historical property insurance claims records and analyzes the causes of accidents (fire, theft, burst water pipes, etc.), frequency of accidents, and average loss amount to quantify its property risk.

[0056] Understandably, this mechanism of automated real-time acquisition of internal and external data achieves immediacy, comprehensiveness, and structure compared to the traditional method of salespersons manually and discretely collecting customer information. In the healthcare insurance field, this mechanism can, with only the company name provided by the customer, construct a data package containing the company's fundamentals, common industry risks, and historical insurance experience within seconds. This provides the AI ​​model with input information approaching the depth of underwriters' professional investigations, greatly improving the speed and initial credibility of the quote initiation. In the property insurance (fintech) field, this mechanism can quickly outline the physical risk profile of the target company (based on publicly available attributes) and industry risk benchmarks (based on internal historical data), enabling the pre-quotation model to provide a valuable quote range based on similar risk patterns, even in the absence of a detailed asset list provided by the customer. This step is a key technological guarantee for achieving the core user experience and commercial value of "getting a quote quickly with minimal information input," providing high-quality, multi-dimensional real-time data fuel for the accurate inference of the model in the subsequent step S160.

[0057] Step S160: Input the enterprise attribute information and the historical business data into the pre-quotation model, and output the reference quotation information of the target object.

[0058] It's important to note that step S160 is the decision-making and output stage of the pre-quotation process. Its core lies in utilizing the established intelligent pre-quotation model to comprehensively analyze real-time aggregated internal and external data, generating a reference quotation that combines numerical accuracy, business reference value, and persuasive communication. This step directly addresses the problems of traditional quotations that only provide a single amount, lack reference basis, and struggle to build customer trust. In the healthcare insurance field, this means not only predicting the premium range but also providing comparable underwriting cases to substantiate the pricing logic; in the property insurance (fintech) field, it means providing a reasonable quotation range while showcasing the actual coverage of similar risk assets to enhance the credibility of the quotation.

[0059] In some embodiments of the present invention, such as Figure 6 As shown, step S160 includes the following steps: Step S161: Generate a numerical range representing the estimated premium range; Step S162: From the historical policy dataset, match at least one historical underwriting case that is most similar to the input features of the target object; Step S163: The numerical range and the de-identified summary information of the historical underwriting cases are output together as the reference quotation information.

[0060] Specifically, in step S161, the estimated premium range is generated by reasoning about the input features using a pre-quotation model and outputting a predicted range with statistical confidence. The model not only predicts point estimates (such as expected premiums) but also calculates the uncertainty of the prediction, generating a confidence interval (e.g., a 90% confidence interval). In the field of healthcare insurance, for the aforementioned "XX Biotechnology Company," after analyzing its industry, size, age structure, and historical data of similar companies, the model might output "estimated annual premium range of [1.8 million yuan, 2.2 million yuan]". In the field of property insurance (fintech), for "YY Catering Management Co., Ltd.", after assessing its industry risk, operating scale, geographical distribution, and claims records of similar companies, the model might output "estimated annual premium range of [80,000 yuan, 120,000 yuan]". This range reflects the uncertainty of risk, providing sales representatives with flexible negotiation space when communicating with clients.

[0061] Specifically, in step S162, the matching of similar historical insurance cases is based on feature similarity calculation. The system filters policies from the historical policy dataset that belong to the same cluster group as the target object (or have highly similar key features, such as industry or size), and uses methods such as Euclidean distance, cosine similarity, or model-based embedding vector similarity to find the N most similar historical cases. In the field of healthcare insurance, for the aforementioned biotechnology company, the system may match three historical insured companies in the same industry, with a staff size between 250 and 350, and similar employee age structures, with actual annual premiums of RMB 1.85 million, RMB 1.95 million, and RMB 2.1 million, respectively. In the field of property insurance, for the catering company, the system may match five historical insured merchants who are also chain restaurants, located in similar business districts, and with similar area, with actual annual premiums of RMB 85,000, RMB 92,000, RMB 101,000, RMB 115,000, and RMB 120,000, respectively.

[0062] Specifically, in step S163, the output of anonymized summary information is to provide valuable reference while protecting customer privacy. The system will anonymize key information of the matched historical cases (such as hiding company names and specific addresses), retaining only the dimension information and actual premiums that are valuable for risk assessment and pricing. In the field of medical and health insurance, the output summary information may be: "Reference Case 1: A technology company in the same industry, with approximately 280 employees, an average age of 32, and a premium of RMB 1.95 million last year; Reference Case 2: A biotechnology company in the same industry, with approximately 320 employees, an average age of 35, and a premium of RMB 2.1 million last year." In the field of property insurance, the output summary information may be: "Reference Case 1: A chain Chinese restaurant in the same business district, with an area of ​​200 square meters, an insured property value of RMB 5 million, and a premium of RMB 92,000 last year; Reference Case 2: A fast food restaurant of the same type, with an area of ​​180 square meters, an insured property value of RMB 4.5 million, and a premium of RMB 85,000 last year."

[0063] Understandably, this dual-track output mechanism of "numerical range + similar cases" greatly enriches the dimensions and value of pricing information compared to the traditional method of only outputting a single price figure. In the healthcare insurance field, this mechanism not only provides a scientific pricing range but also offers strong market evidence and justification for the price quote through real and comparable underwriting cases in the same industry. This helps salespeople explain the basis of their quotes more professionally and confidently when communicating with corporate human resources departments or purchasing decision-makers, accelerating the decision-making process. In the property insurance (fintech) field, this mechanism allows salespeople to intuitively show clients "what is the typical market cost of protection for companies similar to yours." This transparent reference greatly enhances the objectivity and persuasiveness of the quote, especially suitable for SME clients with limited insurance knowledge, helping to reduce communication costs and build trust. This step is a key interface design that effectively transforms the technical capabilities of AI models into practical tools for frontline business operations. It not only outputs a number but also the logic and evidence supporting that number, thereby significantly improving the success rate and customer satisfaction of quoting while improving efficiency.

[0064] In some embodiments of the present invention, such as Figure 6 As shown, after step S160, the method further includes step S170, obtaining the actual policy data generated after the target object is underwritten based on the reference quotation information, and updating the pre-quotation model according to the actual policy data.

[0065] It should be noted that step S170 is a crucial closed-loop link in the pre-quotation system of this invention, enabling self-evolution and continuous optimization. Its core lies in using real business results (final order data) as feedback signals to correct and enhance the pre-quotation model, thereby breaking the limitation of traditional quotation models remaining fixed after training. In the field of healthcare insurance, this means feeding back the details of the group insurance policies actually underwritten by corporate clients and their subsequent medical claims experience to the model to calibrate its assessment of similar corporate health risks. In the field of property insurance (fintech), it means using the details of the property insurance policies actually taken out by enterprises and their subsequent claims records to correct and enrich the model's understanding of similar property risks.

[0066] In some embodiments of the present invention, such as Figure 8 As shown, step S170 includes the following steps: Step S171: Calculate the error between the reference quotation information and the actual premium recorded in the real policy data; Step S172: Add the real policy data as a new data sample to the corresponding historical policy data subset; Step S173: Retrain the pre-quotation model based on the updated historical policy dataset.

[0067] Specifically, in step S171, the error is calculated by comparing the estimated range or point value of the pre-quoted price with the actual premium. Various error measurement methods can be used, such as calculating the absolute or relative error between the actual premium and the midpoint of the predicted range. In the field of healthcare insurance, if the pre-quoted price range for "XX Biotechnology Company" is [1.8 million yuan, 2.2 million yuan] (midpoint 2 million yuan), and the final negotiated annual premium is 1.9 million yuan, then the absolute error is 100,000 yuan, and the relative error is 5.26% (10 / 190). The system will record this error. In the field of property insurance (fintech), if the pre-quoted price range for "YY Catering Company" is [80,000 yuan, 120,000 yuan] (midpoint 100,000 yuan), and the actual premium is 95,000 yuan, then the absolute error is 5,000 yuan, and the relative error is approximately 5.26%. This error calculation helps quantify the model's prediction bias for this type of risk.

[0068] Specifically, in step S172, the method for adding real policy data to the historical dataset is to perform data standardization and grouping. First, the final terms information of the real policies (such as the accurate insured amount, deductible, special provisions, and actual premium) is integrated with the attribute information of the target object collected in step S150 and subsequent claims records (generated over time) to form a complete and uniformly formatted new data sample. Then, the system uses the clustering model or rules established in step S120 to automatically classify this new sample into its policy group (such as the "medium data volume - pharmaceutical manufacturing - 300-person scale" group) and add the sample to the corresponding historical policy data subset of that group. In the field of health insurance, the aforementioned biotechnology company's actual insured policy of 1.9 million yuan, along with its employee structure and subsequent possible claims records, is classified as a new sample into the "pharmaceutical manufacturing" related group. In the field of property insurance, the aforementioned catering company's actual insured policy of 95,000 yuan, along with its store details and subsequent possible claims events, is classified into the "chain catering" related group.

[0069] Specifically, in step S173, the retraining of the pre-quotation model can be based on a triggering strategy (such as accumulating a certain number of new samples) or a periodic strategy (such as monthly / quarterly). Retraining does not involve a full reconstruction each time; instead, it can involve incremental learning or periodic full updates. For a specific policy group that receives new samples, the system will use the updated data subset to re-execute steps S130 to S140: that is, re-select / derive features based on the new data distribution, and retrain the corresponding prediction sub-model using the pre-set training method for that group (hold-out method for high data volume, cross-validation for low data volume). In the healthcare insurance field, when the "Internet technology companies" group accumulates enough new underwriting samples, the system will trigger the retraining of the sub-model for that group, enabling the model to learn the latest protection trends and health risk changes in the industry. In the property insurance field, as data for emerging high-risk groups such as "new energy vehicle charging stations" gradually becomes richer, the prediction accuracy of the corresponding sub-model will be significantly improved through retraining.

[0070] Understandably, through this closed-loop learning and iterative update mechanism based on real business feedback, the pre-quotation system of this invention achieves a qualitative leap from a "static model" to a "self-evolving intelligent agent." In the field of healthcare insurance, this mechanism enables the model to keep pace with medical inflation, changes in disease patterns, and adjustments in corporate welfare policies, dynamically calibrating health risk pricing benchmarks for different industries and companies of different sizes, ensuring that quotation recommendations always align with current market risk levels. In the field of property insurance (fintech), this mechanism enables the model to promptly absorb underwriting experience from emerging risk sectors (such as the sharing economy and new manufacturing industries), as well as changes in existing risk areas' payout models, continuously improving its ability to assess long-tail and complex risks. This closed-loop optimization not only continuously improves the accuracy of individual quotations, but more importantly, it endows the system with long-term vitality and adaptability, solving the industry problem of traditional actuarial models or machine learning models becoming rapidly outdated due to data lag. As the system's operating time increases and the data flywheel spins, the accuracy and reliability of the pre-quotation model will spiral upwards, building for the company a dynamic pricing intellectual asset based on real business data that is difficult for competitors to replicate.

[0071] In summary, the solution implemented in this embodiment of the invention can solve the problems of insufficient model generalization ability, insufficient credibility of pricing results, low model training efficiency, and lack of continuous self-optimization ability in existing insurance pre-quotation methods.

[0072] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0073] In one embodiment, a data mining-based insurance product pre-quotation device is provided, which corresponds one-to-one with the data mining-based insurance product pre-quotation method described in the above embodiments. For example... Figure 9 As shown, the data mining-based insurance product pre-quotation device includes a first acquisition module 910, an analysis module 920, a screening module 930, a model building module 940, a second acquisition module 950, and an output module 960. Detailed descriptions of each functional module are as follows: The first acquisition module 910 is used to acquire a historical policy dataset, which includes the dimensional feature data, premium data and claims data of the policies; Analysis module 920 is used to perform cluster analysis on the historical policy dataset using a clustering algorithm to obtain multiple policy groups with different risk characteristics; The filtering module 930 is used to filter out a set of features that are correlated with the premium data and meet preset conditions based on the results of the clustering analysis from the dimensional feature data and the claims data. The model building module 940 is used to build a pre-quotation model based on the feature set; The second acquisition module 950 is used to acquire the enterprise attribute information of the target object and its associated historical business data in response to a pre-quotation request for the target object; The output module 960 is used to input the enterprise attribute information and the historical business data into the pre-quotation model and output the reference quotation information of the target object.

[0074] In one embodiment, the analysis module 920 is specifically used for: Extract multiple preset dimensional features from the historical policy dataset; Unsupervised clustering algorithms are used to process the extracted dimensional features to aggregate insurance policies with similar risk characteristics into the same group. Count the number of policies contained in each group, and divide the group into high data volume group and low data volume group based on the number.

[0075] In one embodiment, the screening module 930 is specifically used for: For each policy group, calculate the correlation coefficient between each feature and the premium data. Features with correlation coefficients greater than a first threshold are selected to form an initial feature subset; At least two features in the initial feature subset are combined to generate at least one derived feature, and the derived feature is added to the feature set.

[0076] In one embodiment, the model building module 940 is specifically used for: For data that is divided into high-volume single groups, the hold-out method is used to divide the training set and the validation set to train the first prediction sub-model. For data that is divided into low-volume single groups, cross-validation is used to train the second prediction sub-model. The first prediction sub-model and the second prediction sub-model are integrated to form the pre-quotation model.

[0077] In one embodiment, the second acquisition module 950 is specifically used for: By calling an external enterprise information query interface, the publicly available enterprise information of the target object can be obtained; Based on at least one key attribute in the publicly available corporate information, retrieve historical underwriting records with similar attributes from the internal underwriting database; Based on at least one key attribute in the publicly available corporate information, retrieve relevant historical claims records from the internal claims database.

[0078] In one embodiment, the output module 960 is specifically used for: Generate a numerical range representing the estimated premium range; From the historical policy dataset, match at least one historical underwriting case that is most similar to the input features of the target object; The numerical range and the anonymized summary information of the historical underwriting cases are combined and output as the reference quotation information.

[0079] In one embodiment, the output module 960 is further configured to: Calculate the error between the reference price information and the actual premium recorded in the actual policy data; The actual policy data is used as a new data sample and added to the corresponding historical policy data subset; The pre-quote model is retrained based on the updated historical policy dataset.

[0080] This invention provides a solution for an insurance product pre-quotation device based on data mining. Through data-driven historical policy clustering analysis and differentiated modeling, policies are grouped according to multi-dimensional characteristics such as industry and scale, and the risk-premium relationship is extracted. This effectively overcomes the insufficient generalization ability caused by the flattening of data in traditional methods. By employing differentiated verification strategies and feature engineering for high and low data volume groups, the model's stability is ensured in data-scarce scenarios while fully exploring potential patterns in rich data scenarios, significantly improving the adaptability and prediction accuracy of the pre-quotation model for different customer groups. This solution enables sales agents to quickly and independently obtain relatively accurate reference quotes under conditions of limited customer information and scarce underwriting resources, technically solving the core contradictions of slow response, high information barriers, and missed business opportunities in traditional quotation processes.

[0081] Furthermore, the core reason for the insufficient credibility and practicality of pre-quotation results in related technologies is the isolated output information and lack of business context. This solution, however, while generating a reference quotation range, simultaneously queries external company information and internal underwriting and claims databases, automatically matching and presenting similar historical underwriting cases that have undergone anonymization. This places the pre-quotation results within a real business reference system, providing sales personnel with an intuitive and credible basis for communication. This not only improves the acceptability of quotations but also enhances the professional image and responsiveness of the sales team, fundamentally improving the practicality and persuasiveness of the pre-quotation system in real sales scenarios.

[0082] In summary, the above-mentioned data mining-based insurance product pre-pricing method can solve the problems of insufficient model generalization ability, insufficient credibility of pricing results, and low model training efficiency in existing insurance pre-pricing methods.

[0083] Specific limitations regarding the data mining-based insurance product pre-quotation device can be found in the limitations of the data mining-based insurance product pre-quotation method described above, and will not be repeated here. Each module in the aforementioned data mining-based insurance product pre-quotation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0084] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements a recommended method for an optimal strategy, representing the functions or steps on the server side.

[0085] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a recommended method for an optimal strategy, representing client-side functions or steps.

[0086] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain historical policy datasets, which include policy dimensional feature data, premium data, and claims data; Clustering algorithms were used to perform cluster analysis on the historical policy dataset to obtain multiple policy groups with distinct risk characteristics; Based on the results of the cluster analysis, a set of features that are correlated with the premium data and meet the preset conditions are selected from the dimensional feature data and the claims data. Based on the aforementioned feature set, a pre-quotation model is constructed; In response to a pre-quotation request for a target object, obtain the target object's enterprise attribute information and its associated historical business data; The enterprise attribute information and the historical business data are input into the pre-quotation model, and the reference quotation information of the target object is output.

[0087] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain historical policy datasets, which include policy dimensional feature data, premium data, and claims data; Clustering algorithms were used to perform cluster analysis on the historical policy dataset to obtain multiple policy groups with distinct risk characteristics; Based on the results of the cluster analysis, a set of features that are correlated with the premium data and meet the preset conditions are selected from the dimensional feature data and the claims data. Based on the aforementioned feature set, a pre-quotation model is constructed; In response to a pre-quotation request for a target object, obtain the target object's enterprise attribute information and its associated historical business data; The enterprise attribute information and the historical business data are input into the pre-quotation model, and the reference quotation information of the target object is output.

[0088] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0091] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for pre-quoting insurance products based on data mining, characterized in that, include: Obtain historical policy datasets, which include policy dimensional feature data, premium data, and claims data; Clustering algorithms were used to perform cluster analysis on the historical policy dataset to obtain multiple policy groups with distinct risk characteristics; Based on the results of the cluster analysis, a set of features that are correlated with the premium data and meet the preset conditions are selected from the dimensional feature data and the claims data. Based on the aforementioned feature set, a pre-quotation model is constructed; In response to a pre-quotation request for a target object, obtain the target object's enterprise attribute information and its associated historical business data; The enterprise attribute information and the historical business data are input into the pre-quotation model, and the reference quotation information of the target object is output.

2. The insurance product pre-pricing method based on data mining according to claim 1, characterized in that, The historical policy dataset is clustered using a clustering algorithm to obtain multiple policy groups with distinct risk characteristics, including: Extract multiple preset dimensional features from the historical policy dataset; Unsupervised clustering algorithms are used to process the extracted dimensional features to aggregate insurance policies with similar risk characteristics into the same group. Count the number of policies contained in each group, and divide the group into high data volume group and low data volume group based on the number.

3. The insurance product pre-pricing method based on data mining according to claim 1, characterized in that, Based on the results of the cluster analysis, a set of features that meet preset conditions and are correlated with the premium data is selected, including: For each policy group, calculate the correlation coefficient between each feature and the premium data. Features with correlation coefficients greater than a first threshold are selected to form an initial feature subset; At least two features in the initial feature subset are combined to generate at least one derived feature, and the derived feature is added to the feature set.

4. The method for pre-quoting insurance products based on data mining according to claim 2, characterized in that, The step of constructing a pre-quotation model based on the feature set includes: For data that is divided into high-volume single groups, the hold-out method is used to divide the training set and the validation set to train the first prediction sub-model. For data that is divided into low-volume single groups, cross-validation is used to train the second prediction sub-model. The first prediction sub-model and the second prediction sub-model are integrated to form the pre-quotation model.

5. The method for pre-quoting insurance products based on data mining according to claim 1, characterized in that, The acquisition of the target object's enterprise attribute information and its associated historical business data includes: By calling an external enterprise information query interface, the publicly available enterprise information of the target object can be obtained; Based on at least one key attribute in the publicly available corporate information, retrieve historical underwriting records with similar attributes from the internal underwriting database; Based on at least one key attribute in the publicly available corporate information, retrieve relevant historical claims records from the internal claims database.

6. The method for pre-quoting insurance products based on data mining according to claim 1, characterized in that, The output of reference quotation information for the target object includes: Generate a numerical range representing the estimated premium range; From the historical policy dataset, match at least one historical underwriting case that is most similar to the input features of the target object; The numerical range and the anonymized summary information of the historical underwriting cases are combined and output as the reference quotation information.

7. The method for pre-quoting insurance products based on data mining according to claim 1, characterized in that, After inputting the enterprise attribute information and the historical business data into the pre-quotation model and outputting the reference quotation information of the target object, the method further includes: Obtain the actual policy data generated after the target object is underwritten based on the reference quotation information, and update the pre-quotation model according to the actual policy data; Specifically, updating the pre-quote model based on the actual policy data includes: Calculate the error between the reference price information and the actual premium recorded in the actual policy data; The actual policy data is used as a new data sample and added to the corresponding historical policy data subset; The pre-quote model is retrained based on the updated historical policy dataset.

8. A data mining-based insurance product pre-quotation device, characterized in that, include: The first acquisition module is used to acquire historical policy datasets, which include policy dimensional feature data, premium data, and claims data. The analysis module is used to perform cluster analysis on the historical policy dataset using a clustering algorithm to obtain multiple policy groups with different risk characteristics; The filtering module is used to filter out a set of features that are correlated with the premium data and meet preset conditions based on the results of the clustering analysis. The model building module is used to build a pre-quotation model based on the feature set; The second acquisition module is used to acquire the enterprise attribute information of the target object and its associated historical business data in response to a pre-quotation request for the target object; The output module is used to input the enterprise attribute information and the historical business data into the pre-quotation model and output the reference quotation information of the target object.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the data mining-based insurance product pre-quotation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the data mining-based insurance product pre-quotation method as described in any one of claims 1 to 7.