Intelligent insurance pricing method based on public network information

By obtaining data from public network information and building a three-level database, and combining it with internal data to build a weighted pricing model, we have solved the problems of limited data sources and delayed updates in traditional insurance pricing, improved the scientificity and accuracy of the pricing model, reduced costs, and supported the development of innovative products and rapid response.

CN120634657AActive Publication Date: 2025-09-12上海济物光电技术有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511128523.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In the traditional insurance pricing model, data sources are limited, updates are delayed, costs are high, and there is a lack of heterogeneous data integration mechanisms, resulting in insufficient accuracy in the pricing model and an inability to meet the risk measurement needs of innovative products.

Method used

By building multi-source heterogeneous data retrieval logic, obtaining data from public network information, establishing a three-level database for data calibration and governance, building a weighted pricing model based on internal data, and monitoring data source changes and model effects in real time, automatic optimization can be achieved.

Benefits of technology

It expands the scope of pricing factors, improves the scientificity and accuracy of pricing models, reduces costs, supports innovative product development, and realizes dynamic optimization and rapid response of models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634657A_ABST
    Figure CN120634657A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent insurance pricing method based on public network information, and belongs to the technical field of information processing. According to the method, a public network information-based intelligent insurance pricing method architecture and three-level databases (an original database, a calibration database and an application database) are constructed, a multi-source heterogeneous data retrieval logic, a data availability multi-dimensional calibration rule and a treatment method are designed, and a public network information and intranet information combined weight type insurance pricing model is established. And the whole process of automatic acquisition, storage, calibration, treatment, modeling and optimization of public network information is realized. According to the method, the scope of insurance pricing factors is expanded, the data granularity and the scientificity and accuracy of the pricing model are improved, the development cost is reduced, and the method is suitable for dynamic pricing requirements of innovative and traditional insurance products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information processing technology, and in particular relates to an insurance intelligent pricing method based on public network information. Background Art

[0002] In the insurance industry, product pricing hinges on the accuracy of actuarial models, which rely on sufficient and reliable risk factor data. Traditional pricing models rely primarily on data from insurance companies' internal databases or third-party partner organizations. This data is typically standardized and can be used directly in modeling, but it has the following limitations:

[0003] Limited data sources: For innovative insurance products, there is a lack of historical data or similar product references, making it difficult to meet the risk measurement needs of new areas; for traditional products, there are also problems such as incomplete coverage of internal database factors and insufficient data, which affect the accuracy of the pricing model.

[0004] Update lag: Internal data relies on manual collection and third-party integration, with a long update cycle (usually six months or a year), which cannot adapt to the rapid changes in product application scenarios and the insurance product market.

[0005] High cost: Data procurement, platform integration and manual modeling consume a lot of time and money, creating a high barrier to entry, especially for small and medium-sized insurance companies.

[0006] Data availability issue: In the insurance business, intranet data is generally considered reliable and can be directly used for pricing modeling, but there is a lack of data availability assessment technology.

[0007] Data heterogeneity problem: Unstructured data also contains information related to insurance business risks, but the processing of heterogeneous data is currently limited to basic aggregation, association and query. Insurance product pricing is also limited to the use of structured data, and there is a lack of effective heterogeneous data integration mechanism.

[0008] Public online information (such as government websites, research sites, and professional and social media) offers extensive coverage, rapid updates, and low access costs. It represents a valuable, untapped data resource for insurance product pricing. However, public online information is significantly multi-sourced and heterogeneous, resulting in varying degrees of availability. Key challenges, such as assessing data availability and integrating heterogeneous data, must be addressed before it can be applied to insurance product pricing modeling.

[0009] Some existing technologies attempt to supplement internal databases with single external network data, but these efforts fail to address the availability calibration, data governance, application modeling, and automated model updates required for incorporating public network information into insurance product pricing. Therefore, there is an urgent need for an insurance pricing technology that can efficiently integrate multi-source, heterogeneous public network information, enabling intelligent calibration and dynamic optimization. This will improve the scientificity and accuracy of models, reduce costs, and support innovative product development. Summary of the Invention

[0010] To solve the above technical problems, the present invention provides an intelligent insurance pricing method based on public network information, which uses public network information to expand product pricing factors, increase data granularity, and give full play to the advantages of the public network in terms of the breadth, convenience and cost of obtaining data, to meet the needs of timely iterative updates of models, innovative product development and reducing data acquisition costs.

[0011] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0012] An intelligent insurance pricing method based on public network information, the method comprising:

[0013] Step 1: Public network information data retrieval: Based on insurance product types and risk factor requirements, build multi-source heterogeneous data retrieval logic and obtain raw data through the public network;

[0014] Step 2: Design of a three-level public network information database: Build a level 0 database (original database) to store the original data collected from the public network; build a level 1 database (calibration database) to calibrate the availability of the original data; build a level 2 database (application database) to manage and store the calibration database data;

[0015] Step 3: Collaborative pricing modeling: Combine data from public information application databases with varying availability levels with data from insurance companies' internal databases to build a weighted insurance product pricing model weighted by data availability level.

[0016] Step 4: Automatic model optimization: Real-time dynamic monitoring of changes in public network data sources and the application effect of weighted insurance product pricing models. Automatically update the pricing model through a feedback mechanism, adjust data calibration and model weighting rules, and achieve closed-loop optimization and continuous iteration of the pricing model.

[0017] The beneficial effects of the present invention are:

[0018] Expanding the scope of pricing factors: Compared to internal data platforms, public network data has the characteristic of being widely sourced and unrestricted. By collecting and applying public network information, this invention expands the scope of factors available for pricing models and adds new risk factor dimensions, helping to improve the scientific nature and accuracy of pricing models. It also effectively addresses the problem that internal data cannot meet the risk calculation requirements of insurance product innovation.

[0019] Improving the granularity of forecast data: This invention expands the sample size of factor data for existing pricing models through public networks, increasing granularity based on intranet data, and helping to improve pricing model accuracy. Leveraging public data intelligent proxy services enables high-frequency data acquisition, aligning with the frequency of public data releases, and compensating for the shortcomings of internal data in application scale and time update frequency. Furthermore, high-frequency public data improves the ability to configure forecast scales for different time observations during the modeling process.

[0020] Reduced pricing model development and application costs: This method enables automated online iterative optimization through automated model result monitoring and comparison, as well as backtracking and feedback on constraint conditions, reducing the manpower and time costs of model development and application. Furthermore, by fully leveraging public data, it effectively reduces the capital costs associated with purchasing and integrating data, compared to building internal data platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a system architecture diagram of the insurance intelligent pricing method based on public network information of the present invention;

[0022] Figure 2 This is a technical flow chart of the insurance intelligent pricing method based on public network information of the present invention;

[0023] Figure 3 Schematic diagram of the public network information level 0 database (original database) of the present invention;

[0024] Figure 4 Schematic diagram of the public network information level 1 database (calibration database) of the present invention;

[0025] Figure 5 Schematic diagram of the public network information level 2 database (application database) of the present invention;

[0026] Figure 6 A flowchart of the collaborative pricing modeling technology of the present invention;

[0027] Figure 7 This is a flowchart of the model automatic optimization technology of the present invention, wherein (a) is a schematic diagram of the principle of automatic monitoring of data source changes and iterative updates, and (b) is a schematic diagram of the principle of automatic monitoring of model application effects and backtracking optimization. DETAILED DESCRIPTION

[0028] The present invention will be further described below with reference to the accompanying drawings and examples.

[0029] This invention leverages AI and big data applications to achieve breakthroughs from zero to one for innovative products lacking internal network information, using public network information. This allows for risk calculation, product solutions, and diversified protection for businesses and individuals. For traditional products, public network information is used to optimize modeling, improving the scientific nature and accuracy of model calculations. Furthermore, automated modeling systems save manpower, time, and financial costs.

[0030] like Figure 1 As shown, dual-network service nodes (public and internal) are established to connect the internal and external network environments. Public network data is accessed on demand, comprised of multiple public data intelligent proxy services deployed on the nodes. After data integration is achieved, the pricing modeling platform is linked to the data source interfaces, enabling interactive model rehearsals. To ensure the security of external public network information data, internal and external data are isolated using technologies such as interface gateways. Public network data is stored in a three-tiered, independently constructed database, with a recursive, application-interactive structure across the three stages of retrieval, processing, and modeling to prevent data contamination and confusion. During the pricing modeling process, the two interfaces work together to generate processed data and store the final output model results. Furthermore, write protection constraints are implemented throughout the entire architecture. When external or internal data is missing or there are application issues, only internal or external data is used for model calculations.

[0031] like Figure 2 As shown, a full-process rule is designed around the public network information database to conduct a multi-dimensional data availability assessment on indicators such as the quality of the acquired public network information. Characteristics such as the relevance of information from different sources are analyzed, and weighted scoring rules are formulated to support the rapid and accurate completion of information collection and risk modeling for insurance pricing. On this basis, based on the actual application effect of the model and external feedback, the front-end logic is fed back to gradually update and improve the evaluation rules for various information sources and types. Specifically, the detailed steps of the embodiment of the insurance intelligent pricing method based on public network information of the present invention include:

[0032] Step 1: Collecting public network information data: Build multi-source heterogeneous data retrieval logic based on insurance product types and risk factor requirements. Through the integrated application of distributed crawlers and search engine APIs, collect structured data from the public network, as well as unstructured and semi-structured data such as text, images, audio, and video. Specifically, this includes:

[0033] Step 1.1: Build search logic;

[0034] Determine a product classification system by insurance type (major, sub-category, and line). Based on product solutions, market demand, and in conjunction with internal databases, identify required risk factors and search keywords, design metadata for public online information retrieval, and optimize indexing of public online information. Develop a product model framework for the required risk factors, establish initial information retrieval logic, and record required elements such as product categories, names, core factors, and search affixes. Categorize each product library by product category and name, and associate relevant search factors for continuous dynamic updates and supplementation.

[0035] Based on the above logic, the index-related information retrieval path and retrieval target are clarified to form a "metadata" index, which is gradually expanded through product demand and category classification storage.

[0036] Taking corporate property insurance as an example, it is categorized into property insurance - non-auto insurance - property damage - corporate property insurance. Based on its risk factors, information such as fire (keywords "fire", "combustion", "firefighting"...) and extreme climate (keywords "heavy rain", "typhoon", "warning"...) can be retrieved.

[0037] Taking agricultural meteorological index insurance as an example, it is categorized into property insurance - agricultural insurance - index type - agricultural meteorological index insurance. Based on its risk factors, information such as temperature (keywords "temperature", "high temperature"...), rainfall (keywords "rainfall", "precipitation"...) and so on are retrieved.

[0038] Step 1.2: Multi-search engine intelligent integration application;

[0039] Compared to internal data integration platforms, public network data comes from a variety of sources and structures. To leverage its advantages in data volume and dimensionality, we intelligently integrate multiple search engines to implement differentiated retrieval logic for numerical values, text, images, audio, and video, capturing and recording the full amount of data required.

[0040] By providing an extensible and adaptive data interface that accommodates multiple public network information sources, we ensure the breadth of collected data as much as possible. At the same time, through interface constraints, we confirm the scope of factors such as the time and type of required information to improve the accuracy of collected data and ensure data application efficiency.

[0041] Tool integration: Through traditional data platforms, various search engine APIs are integrated, and by deploying distributed crawlers, a public information network data framework is formed to expand coverage.

[0042] Conditional association: Associating information retrieval logic, based on product requirements, clarifying search objects, categories and other constraints, and formulating a basic operation framework.

[0043] Intelligent scheduling: Combine search keywords, set up a task scheduler, and dynamically allocate real-time applications.

[0044] Step 2: Design of the three-level public network information database:

[0045] Build a Level 0 database (raw database) to store raw data collected from the public network. Build a Level 1 database (calibration database) to perform multi-dimensional usability calibration on the raw data based on quality, integrity, and information volume. Build a Level 2 database (application database) to store insurance pricing data that has been cleaned, integrated, mined, and governed. Specifically, it includes:

[0046] Step 2.1: Build the level 0 database (original database);

[0047] For existing products, in order to integrate with the internal data platform, the service interface is expanded on the basis of the previous platform to form a public data interface and differentiated groups to store and retrieve information for application; for innovative products, the level 0 database is independently grouped and applied through an independent interface.

[0048] The public network data stored in the Level 0 database includes both structured data such as numerical values ​​and unstructured or semi-structured data such as text, images, audio, and video. It is allowed to aggregate the data of various risk factors and build a comprehensive Level 0 database (see the attached Figure 3 ), or for each risk factor data, a level 0 sub-database of a single factor is constructed, and the collection of all level 0 sub-databases is used as the level 0 database.

[0049] Step 2.2: Design of data availability calibration rules and construction of level 1 database (calibration database);

[0050] In the current pricing model, insurance companies rely heavily on internal data from third-party platforms, allowing them to submit data suggestions only through communication. This lacks data evaluation rules and prevents them from independently evaluating data based on front-end business needs. Consequently, during the pricing process, internal data is often assumed to be reliable, modeled uniformly, or manually adjusted based on subjective judgment.

[0051] Step 2.2.1, Data availability calibration rule design;

[0052] Public network information comes from a wide range of sources, has diverse structures, and has varying usability. Therefore, it must be calibrated and governed before it can be used in pricing models. Therefore, this paper designs data calibration rules for raw data obtained from the public network. Based on pre-set rules, the raw data is calibrated across multiple dimensions of data usability, including information quality, information integrity, and information volume.

[0053] The data calibration rules are set as follows:

[0054] ① Information Quality Level: This evaluates the reliability of data, with four levels: "Reliable," "Relatively Reliable," "Basically Reliable," and "Unreliable." This is based on the reliability of the data source. This includes: 1) the authority of the information source and its dissemination path; 2) the logical rationality of the information content, the integrity of the chain of evidence, and the identification of manipulation; and 3) the transparency of the information platform, content oversight, and industry relevance. For example: government official website releases > business organization releases > research institution releases > self-media releases.

[0055] ② Information integrity level: Evaluates the matching level of data coverage with application requirements, with five levels: 5, 4, 3, 2, and 1. The information integrity level is related to application requirements and is calculated as: Collection data coverage / Application required data coverage Rounding is done to the nearest integer. For example, if the model requires the average monthly temperature for the entire year, if there are 11 monthly average temperatures, the integrity level is 5; if there are 6 monthly average temperatures, the integrity level is 3; if there are 2 monthly average temperatures, the integrity level is 1.

[0056] ③ Information Level: Evaluates the depth of application information the data can provide. This level is divided into three categories: "quantitative," "semi-quantitative," and "qualitative." For example, information that contains specific numerical values ​​is considered quantitative, information that contains the range of variation in the data is considered semi-quantitative, and information that contains qualitative descriptions is considered qualitative.

[0057] Data calibration rules can be further adjusted and improved based on the gradual enrichment of public network information and the expansion of application needs.

[0058] Step 2.2.2, build the level 1 database (calibration database);

[0059] According to the "data calibration rules", the availability of the level 0 database data is evaluated from the above dimensions (which can be further expanded), and each data and dimension is marked and stored to form a level 1 database (refer to Figure 4 );

[0060] During implementation, the backend service automatically associates pre-written calibration rules with the original database, assesses the quality, integrity, and volume of public network data, and then assigns labels to the data. This automated storage creates a calibration database for subsequent use.

[0061] Step 2.3: Establish data governance methods and build a second-level database (application database);

[0062] Step 2.3.1: Establish data governance methods;

[0063] Establish a public network information data governance method to enable data to be used for insurance pricing modeling. The details are as follows:

[0064] For unstructured and semi-structured data such as text, images, audio and video, we integrate intelligent tools such as NLP technology, computer vision and speech recognition to extract the required information and convert it into structured data.

[0065] For calibration data with semi-quantitative and qualitative information, we conduct information mining with quantification as the goal, using methods such as information mapping, machine learning, and deep learning to increase the information content of each data set. The information quality and integrity levels remain unchanged. For example, based on historical statistical data on the relationship between water quality grades in a particular area and specific water quality index values, we can deduce the water quality index values ​​from the water quality grades.

[0066] For calibration data with insufficient information integrity, with the goal of fully matching application requirements, the relevant incomplete data groups are fused at the "data level" (decision-level fusion is not used to avoid information loss) through data processing (combination, extrapolation and interpolation), AI comprehensive analysis (Kalman filtering, neural network), etc. The information quality level is the weighted average of the information quality levels of all data groups involved according to the information integrity level.

[0067] Threshold conditions for the number of independent final information sources and the relative deviation of the data are set. For each group of calibration data that meets the set source number and deviation threshold conditions and has the same information quality level (except for the unreliable level), the information quality level is improved by one.

[0068] Based on the above governance, data cleaning was performed, successively eliminating data with semi-quantitative and qualitative information levels, integrity levels less than 3, and unreliable quality levels. The remaining data was later labeled with availability levels and stored.

[0069] Step 2.3.2: Build the second-level database (application database);

[0070] Apply data governance methods to perform data format conversion, information mining, information data level fusion, information quality calibration upgrade and data cleaning on the level 1 database data to improve data availability, generate structured data with availability level tags with quality level of basic reliability or above, integrity level of 3 or above, and quantitative information level, and store them to form a level 2 database (such as Figure 5 shown).

[0071] Step 3: Collaborative pricing modeling:

[0072] The data of the public network information level 2 database (application database) with different availability levels is integrated with the data of the insurance company's internal information database, which is considered to be reliable, to build a weighted insurance product pricing model that is weighted by data availability level. Figure 6 shown):

[0073] 1) For the public network information application database data, establish a weight assignment rule based on its final data availability level. The higher the level, the greater the value. Specifically:

[0074] The quality level "reliable" has a weight of 1, "relatively reliable" has a weight of 0.8, and "basic reliable" has a weight of 0.6;

[0075] Integrity level "5" has a weight of 1, level "4" has a weight of 0.8, and level "3" has a weight of 0.6.

[0076] After the information is mined and cleaned, what is retained is the "quantitative" level data, which has a weight of 1.

[0077] The data quality weight is multiplied by the integrity weight and the information weight to obtain the final comprehensive weight of sample data availability. Based on the feedback from the model application effect, it is decided whether the assignment rules need to be adjusted.

[0078] 2) The public network application database is linked to the internal database to form an intermediate result data table that combines internal and external data for model pricing application.

[0079] 3) Divide the public network database and internal database data related to the target product into training sets and test sets, use existing dimensional factors as model features, implement modeling operations based on relevant platforms, and calculate and estimate loss risks.

[0080] 4) During the model calculation process, compared with the previous sample (arithmetic mean) risk loss function based on internal data, a weighted risk loss function model adapted to public network data is established:

[0081] ,

[0082] in, is the predicted risk loss of a factor, is the sample loss value of the factor, for The weight of public network data is a comprehensive value assigned to its availability level. When used for intranet and intranet collaborative pricing modeling, the weight of intranet samples is set to 1.

[0083] If the predicted risk loss of a factor by internal data is , including samples; at the same time, the external network added samples, and their sample loss values ​​are , the comprehensive weights of availability are , then the weighted risk loss function model of internal and external network collaboration is:

[0084] ,

[0085] in, It is the predicted risk loss of the internal and external network collaboration of the factor.

[0086] The weights mentioned above specifically refer to the comprehensive weights of the multi-dimensional availability levels of sample data, which are used to calculate the pure risk estimation of the factors and do not involve the insurance company's adjustment of the rate weights based on market and customer factors.

[0087] 5) Calculate the total risk loss function of the insurance product based on the risk loss functions of each factor involved. The specific model form is determined by the insurance product design and can be linear, nonlinear, or covariance matrix, and the present invention can directly apply it.

[0088] 6) For innovative products lacking intranet data, only a public data interface is provided, allowing modeling to proceed directly from the public information application database. For traditional business products, if security issues or other anomalies arise on the public network, the public data interface is disconnected and pricing modeling based on intranet data is reverted. In either case, it can be considered a special case of the weighted pricing model.

[0089] Taking agricultural meteorological index insurance as an example, previously, internal database risk factors included crop humidity and temperature. However, through external public network information collection, environmental pollution and waterlogging factors can now be incorporated.

[0090] Step 4: Automatic model optimization:

[0091] Traditional insurance product pricing relies solely on fixed model results. This modeling process is manually repeated every six months or annually, incorporating new data within the legacy framework and adjusting factors through manual judgment. Consequently, this traditional model iteration and update process is slow, unable to meet the market's demand for rapid response, and requiring significant labor and time.

[0092] The present invention monitors the changes in front-end data sources and model output effects in real time for the established weighted insurance product pricing model that is weighted by data availability level, and automatically implements continuous iteration and closed-loop optimization in the system background.

[0093] Specifically, there are two types of approaches (such as Figure 7 (a)-(b) of the figure, where (a) is the principle diagram for iterative update of automatic monitoring of data source changes, and (b) is the principle diagram for automatic monitoring of model application effects and backtracking optimization):

[0094] First, automatically monitor data source changes and iteratively update the model;

[0095] 1) The user sets the update frequency (e.g., daily or weekly), and the system automatically refreshes periodically through the public network data interface connected to the backend. Once a new data source is discovered or an existing data source shows data updates, step 1 is initiated to collect data.

[0096] 2) Automatically execute steps 2 to 3 to sequentially update data in public network databases at all levels, continuously optimize the pricing model, and match the latest data status on the front end.

[0097] Second, automatically monitor the model application effect and back-optimize the rules;

[0098] 1) Table the model results, write the retrospective monitoring logic, continuously monitor and compare the model's predicted risk and actual risk trends, and identify the differences.

[0099] 2) For cases where the difference exceeds the business warning limit, extraction and attribution analysis are carried out to locate its deviation characteristics and define the scope of sample data to be optimized.

[0100] 3) Optimize the data availability level assignment logic, simulate the preset model training and output effect comparison to confirm whether the model optimization results meet business requirements.

[0101] 4) If the standards are not met, backtrack and optimize the availability level improvement and cleansing logic associated with the data governance process, simulate the preset model training and output effect comparison to confirm whether the model optimization results meet business requirements.

[0102] 5) If the standards are not met, backtrack and optimize the usability rule logic associated with the data calibration process, simulate the preset model training and output the effect comparison.

[0103] 6) Determine the pricing model that works best given the existing data and factor dimensions.

[0104] 7) Based on actual business needs, add new factor dimensions and expand the front-end data interface to refine the model risk assessment architecture.

[0105] In summary, this invention incorporates public information platform data into pricing modeling and combines it with internal information platform data to expand the application scope of insurance services and enhance their application capabilities. It also fully constructs a comprehensive process architecture encompassing seven key components: data framework, retrieval logic, data collection, data calibration, data governance, collaborative modeling, and model optimization. Compared to traditional insurance product pricing models, this approach enables dynamic updates and optimization, improving model accuracy without the need for manual, repetitive calculations.

[0106] The public information application service provided by this invention can continuously expand data interfaces and incorporate new, multi-source, heterogeneous data sources into the model database. Compared to the fixed internal data docking platform of traditional pricing models, this service offers the advantages of more extensive information and more timely updates. Furthermore, the iterative use of a three-tiered database model for raw data, calibration data, and application data avoids data contamination and enables the integrated application of public and internal data for both innovative and traditional products.

[0107] Compared to traditional product internal data platforms, which typically process and format data in a standardized format, public network data comes from a wide range of sources, from professional business websites and research sites to self-media platforms. These sources include not only the structured intranet data used for traditional insurance product pricing, but also unstructured and semi-structured data such as text, images, audio, and video. This invention integrates intelligent application tools, combines keyword-based filtering logic, and stores multi-source heterogeneous data. It then uses a differentiated scheduling approach to capture all required information, expanding the breadth of data.

[0108] Compared with intranet data, public network data has significant multi-source heterogeneity and mixed characteristics. Data availability must be evaluated and improved in order to make good use of it and give full play to its benefits. The present invention designs data availability level classification rules from multiple dimensions such as information quality, information integrity and information volume, and determines the availability of the collected data based on this. By calibrating the data obtained from the public network, it is possible to directly lock the product factor requirements based on the front-end business feedback to expand the retrieval of data. On this basis, the availability of data is improved through targeted governance methods in each dimension. Through the calibration of the availability of data obtained from the public network and the availability improvement process, the advantages of the public network's wide range and large amount of information are fully utilized, and the accuracy of the pricing model is improved.

[0109] Compared to traditional insurance product pricing models that require manual updates every six months or even a year, this invention implements automatic system detection and, combined with built-in optimization logic, enables real-time iterative optimization of the model. This ensures that the pricing model truly reflects the latest market risk scenarios, enhancing its accuracy and application value.

[0110] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An intelligent insurance pricing method based on public network information, characterized in that: The method comprises: Step 1: Public network information data retrieval: Based on insurance product types and risk factor requirements, build multi-source heterogeneous data retrieval logic and obtain raw data through the public network; Step 2: Design of a three-level public network information database: Construct an original database for storing the original data; construct a calibration database for calibrating the availability of the original data, including: designing data calibration rules, performing multi-dimensional calibration of the original data's information quality level, information integrity level, and information volume level, labeling the data, and forming a calibration database through automated storage. The information quality level is used to evaluate the reliability of the data, the information integrity level is used to evaluate the level of data coverage matching application requirements, and the information volume level is used to evaluate the depth of application information that the data can provide; and construct an application database for data governance and storage of the calibration database data. Step 3: Collaborative pricing modeling: Combine data from public information application databases with varying availability levels with data from insurance companies' internal databases to build a weighted insurance product pricing model weighted by data availability level. Step 4: Automatic model optimization: Real-time dynamic monitoring of changes in public network data sources and the application effect of weighted insurance product pricing models. Automatically update the pricing model through a feedback mechanism, adjust data calibration and model weighting rules, and achieve closed-loop optimization and continuous iteration of the pricing model.

2. The intelligent insurance pricing method based on public network information according to claim 1, characterized in that: Construct an insurance product pricing modeling system that combines public network information with intranet information, specifically including: Build a technical architecture for public network information insurance intelligent pricing applications, providing scalable multi-source heterogeneous data support for insurance product pricing and dual-network services of public network information and internal network information that meet security requirements; Design full-process technology for data collection, data calibration, data governance, data storage, data pricing application modeling and model optimization.

3. The insurance intelligent pricing method based on public network information according to claim 1 is characterized in that: The step 1 includes: establishing a product classification system according to insurance type, setting risk factors and search keywords for each product, and forming a search metadata index; based on the index, collecting structured, unstructured and semi-structured data from the public network through a distributed crawler and search engine API.

4. The insurance intelligent pricing method based on public network information according to claim 1 is characterized in that: The step 2 of constructing an original database for storing the original data collected from the public network includes: For traditional products, public data interfaces and differentiated groups are formed through service interface expansion to store and retrieve information for application. For innovative products, the original database is independently grouped and applied through independent interfaces.

5. The intelligent insurance pricing method based on public network information according to claim 1, characterized in that: In step 2: The information quality level is divided into four levels: "reliable", "relatively reliable", "basic reliable" and "unreliable"; the information integrity level is divided into five levels: 5, 4, 3, 2 and 1; the information quantity level is divided into three levels: "quantitative", "semi-quantitative" and "qualitative".

6. The intelligent insurance pricing method based on public network information according to claim 1, characterized in that: The application database constructed in step 2 for data management and storage of calibration database data includes: For unstructured and semi-structured calibration data, intelligently extract the required information and convert it into structured data; For calibration data with semi-quantitative and qualitative information, we conduct intelligent information mining with quantification as the goal, and improve the information level of each group of data respectively; For calibration data with insufficient information integrity, data-level intelligent fusion is performed on related incomplete data groups with the goal of forming data that fully matches application requirements; Set threshold conditions for the number of independent final information sources and the relative deviation of data. For each set of calibration data that meets the threshold conditions and has the same information quality level, each will be upgraded by one information quality level. Perform data cleaning, re-label and store the retained data according to its availability level to form an application database.

7. The intelligent insurance pricing method based on public network information according to claim 1, characterized in that: The step 3 comprises: For the public network information application database data, the final multi-dimensional data availability level is used to establish a comprehensive availability weight assignment rule. The higher the level, the greater the value. Link the public network information application database with the internal database to form an intermediate result data table that combines internal and external data for use in pricing models; The public network information application database data and internal database data related to the pricing target product are divided into training sets and test sets, and a collaborative pricing model is established to estimate loss risks.

8. The intelligent insurance pricing method based on public network information according to claim 7, characterized in that: Assume that the predicted risk loss of one factor by internal data is , including samples; at the same time, one of the factors mentioned above is added by the public network samples, and their sample loss values ​​are , the comprehensive weights of availability are , then the weighted risk loss function model of internal and external network collaboration is: , in, The predicted risk loss of internal and external network collaboration for one of the factors mentioned above.

9. The intelligent insurance pricing method based on public network information according to claim 8, characterized in that: When only public network data is used, the weighted risk loss function model adapted to public network data is: , in, is the predicted risk loss of a factor, y i is the sample loss value of the factor, i=1,2,…,N, for The comprehensive weight of availability; when used for intranet and intranet collaborative pricing modeling, the intranet sample weight is set to 1.

10. The intelligent insurance pricing method based on public network information according to claim 1, characterized in that: The step 4 comprises: Automatically monitor public network data sources. Once a new data source is discovered, or an existing data source is updated, steps 1-3 are initiated, sequentially updating public network databases at all levels to continuously optimize the pricing model and match the latest front-end data. Automatically monitor the application effect of the pricing model for weighted insurance products. When the difference between the model-predicted risk and the actual risk exceeds the business warning limit, retroactively optimize the data availability level calibration, governance, and weight assignment to improve the model application effect.

Citation Information

Patent Citations

  • Intelligent insurance risk assessment and pricing system based on artificial intelligence

    CN116823496A

  • Insurance product pricing and risk assessment method and device and storage medium

    CN117830009A

  • Intelligent campus operation and maintenance data management method and system based on big data

    CN118503910A

  • Intelligent premium pricing method and system based on multi-dimensional data analysis driving

    CN118887021A

  • Personalized insurance product customization method and risk pricing method based on customer portrait

    CN120163656A