Insurance intelligent pricing method based on public network information

By acquiring data from public network information, establishing a three-level database, and constructing a weighted pricing model, the problems of limited data sources and delayed updates in traditional insurance pricing models are solved. This improves the scientificity and accuracy of the pricing model, supports the development of innovative products, and reduces costs.

CN120634657BActive Publication Date: 2025-12-09上海济物光电技术有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511128523.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-12-09
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Traditional insurance pricing models suffer from limited data sources, delayed updates, high costs, and a lack of heterogeneous data integration mechanisms, resulting in insufficient accuracy of pricing models and difficulty in meeting the risk assessment needs of innovative products.

Method used

By constructing a multi-source heterogeneous data retrieval logic, data is obtained from public network information, a three-level database is established for data labeling and governance, and a weighted pricing model is constructed. Combined with internal data, collaborative pricing is carried out to achieve automatic optimization and iterative updates of the model.

Benefits of technology

It expands the scope of pricing factors, improves the scientific nature and accuracy of pricing models, reduces costs, supports the development of innovative products, and enables dynamic optimization and rapid response of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634657B_ABST
    Figure CN120634657B_ABST
Patent Text Reader

Abstract

The application discloses an insurance intelligent pricing method based on public network information and belongs to the technical field of information processing. The method constructs an insurance intelligent pricing method framework based on public network information, three-level databases (an original database, a calibration database and an application database), designs multi-source heterogeneous data retrieval logic, data availability multi-dimensional calibration rules and management methods, establishes a weight type insurance pricing model combining public network information and internal network information, and realizes the whole process of automatic collection, storage, calibration, management, modeling and optimization of public network information. The application expands the insurance pricing factor category, improves the data granularity, the scientificity and accuracy of the pricing model, reduces the development cost, and is suitable for the dynamic pricing demand of innovative and traditional insurance products.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of information processing, and particularly relates to an insurance intelligent pricing method based on public network information. BACKGROUND

[0002] In the insurance industry, the core of product pricing lies in the accuracy of actuarial models, which relies on sufficient and reliable risk factor data. Traditional pricing models are mainly based on data provided by internal databases of insurance companies or third-party cooperation agencies. These data, which are usually standardized and can be directly used for modeling, have the following limitations:

[0003] Limited data sources: For innovative insurance products, there is a lack of historical data or similar product references, making it difficult to meet the risk measurement needs of new fields; for traditional products, there are also problems such as incomplete factor coverage and insufficient data in internal databases, affecting the accuracy of pricing models.

[0004] Update lag: Internal data relies on manual collection and third-party integration, with a long update cycle (usually half a year or a year), which cannot adapt to the rapid changes in product application scenarios and the insurance product market.

[0005] High cost: Data procurement, platform integration, and manual modeling consume a large amount of time and funds, especially for small and medium-sized insurance companies, forming a high threshold.

[0006] Data availability issues: In insurance business, internal network data is generally considered reliable and can be directly used for pricing modeling, lacking data availability assessment techniques.

[0007] Data heterogeneity issues: Unstructured data also contains insurance business risk-related information, but for heterogeneous data processing, current methods are limited to basic aggregation, correlation, and query, and insurance product pricing is also limited to structured data, lacking effective heterogeneous data integration mechanisms.

[0008] Public network information (such as government websites, scientific research websites, professional and social media) covers a wide range, updates quickly, and is low-cost to obtain. For insurance product pricing, it is a valuable data resource to be developed. However, public network information has significant multi-source heterogeneity characteristics and varying availability. Key issues such as data availability assessment and heterogeneous data integration must be addressed before it can be applied to insurance product pricing modeling.

[0009] In existing technologies, some attempts have been made to supplement internal databases with single external network data, but the issues of availability calibration, data governance, application modeling, and model automatic updating required for applying public network information to insurance product pricing have not been addressed. Therefore, there is an urgent need for an insurance pricing technology that can efficiently integrate multi-source heterogeneous public network information, achieve intelligent calibration and dynamic optimization, to improve the scientificity and accuracy of the model, reduce costs, and support the development of innovative products. SUMMARY

[0010] To solve the above technical problems, the application provides an insurance intelligent pricing method based on public network information, which uses public network information to expand product pricing factors, increases data granularity, and takes advantage of the wide availability, convenience and cost of public networks to meet the needs of timely iteration and updating of models, development of innovative products and reduction of data acquisition costs.

[0011] To achieve the above purpose, the technical scheme adopted by the application is as follows:

[0012] The insurance intelligent pricing method based on public network information comprises the following steps:

[0013] Step 1: Public network information data retrieval: according to the type of insurance products and the demand for risk factors, a multi-source heterogeneous data retrieval logic is constructed, and original data is acquired through a public network;

[0014] Step 2: Three-level public network information database design: a 0-level database (original database) is constructed for storing the original data collected from the public network; a 1-level database (calibration database) is constructed for calibrating the usability of the original data; and a 2-level database (application database) is constructed for data governance and storage of the data in the calibration database;

[0015] Step 3: Collaborative pricing modeling: the public network information application database data with different usability levels are combined with the internal database data of the insurance company to construct a weight-type insurance product pricing model with data usability level weighting;

[0016] Step 4: Model automatic optimization: real-time dynamic monitoring of public network data source changes and weight-type insurance product pricing model application effects, automatic updating of the pricing model, adjustment of data calibration and model weighting rules through a feedback mechanism, and closed-loop optimization and continuous iteration of the pricing model.

[0017] The application has the following beneficial effects:

[0018] Expand the scope of pricing factors: Compared with internal data platforms, public network data has the characteristics of wide and unrestricted sources. The application expands the scope of factors available for the pricing model by collecting and applying public network information, adds new risk factor dimensions, helps to improve the scientificity and accuracy of the pricing model, and effectively solves the problem that internal data cannot meet the risk calculation of insurance product innovation.

[0019] Enhancing the granularity of prediction data: The present application expands the sample size of factor data through public network for the past pricing model, increases the granularity on the basis of internal network data, and helps to improve the accuracy of the pricing model. The intelligent agent service of public data can realize high-frequency data acquisition, align with the public data publishing frequency, make up for the lack of internal data in application scale and time update frequency; at the same time, high-frequency public data improves the ability to configure different time observation prediction scales in the modeling process.

[0020] Reducing the cost of pricing model development and application: The present application realizes online automatic iteration optimization through automatic model result monitoring and comparison, and backtracking of constraint conditions, reduces the manpower and time cost of model development and application. At the same time, the public network data is fully used, compared with the internal data platform, which helps to effectively reduce the fund cost required for purchasing and connecting data. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The system architecture diagram of the insurance intelligent pricing method based on public network information of the present application;

[0022] Figure 2 The technical flow chart of the insurance intelligent pricing method based on public network information of the present application;

[0023] Figure 3 The schematic diagram of the public network information 0-level database (original database) of the present application;

[0024] Figure 4 The schematic diagram of the public network information 1-level database (calibration database) of the present application;

[0025] Figure 5 The schematic diagram of the public network information 2-level database (application database) of the present application;

[0026] Figure 6 The technical flow chart of the collaborative pricing modeling of the present application;

[0027] Figure 7 The technical flow chart of the model automatic optimization of the present application, wherein (a) is the principle diagram of automatic monitoring data source change iteration update, and (b) is the principle diagram of automatic monitoring model application effect and backtracking optimization. DETAILED DESCRIPTION

[0028] The present application will be further described below in combination with the drawings and examples.

[0029] The application realizes the breakthrough from 0 to 1 of the innovative product through the public network information, supports the risk calculation, forms the product scheme, and provides diversified protection for enterprises and individuals. For the traditional product, the public network information is applied to realize the modeling optimization, and the scientificity and accuracy of the model calculation are improved. And the manpower, time, fund and other costs are saved through the automatic model system.

[0030] As shown in Figure 1 , the double network (public network and internal network) service node is built, and the internal and external network environment is connected. The public network data is obtained on demand, and is composed of multiple public data intelligent agent services deployed on the node. After the data integration, the model is realized through the interaction of the interface of each data source. In order to ensure the security of the application of the external public network information data, the interface gate is used to separate the internal and external data. When the public network data is stored, the three-level database is independently constructed, and the data is recursively processed in the three links of retrieval, processing and modeling. In the pricing modeling process, the two interfaces jointly form the processed data and store the final output model result. At the same time, the whole architecture system is written with protection constraints, and when the external data / internal data is missing or there is an application problem, only the internal data / external data is used for model calculation.

[0031] As shown in Figure 2 , around the public network information database, the whole process rule is designed, the quality indicators of the obtained public network information are evaluated in multiple dimensions, the correlation of information from different sources is analyzed, and the weight assignment rule is formulated, so as to support the fast and accurate completion of the information collection and risk modeling of insurance pricing. On this basis, according to the actual application effect of the model and the external feedback, the front-end logic is fed back to gradually update and improve the evaluation rules of various information sources and various types of information. Specifically, the detailed steps of the insurance intelligent pricing method based on the public network information embodiment of the application include:

[0032] Step 1, public network information data collection: according to the type of insurance product and the demand of risk factor, the multi-source heterogeneous data retrieval logic is constructed. Through the integration of distributed crawler and search engine API, the structured data, as well as the unstructured data and semi-structured data such as text, image, audio and video are collected from the public network. Specifically, it includes:

[0033] Step 1.1, retrieval logic construction;

[0034] Determine the product classification system by risk category (major category, subcategory, and line), confirm the required risk factors and search keywords based on product solutions, market demand, and internal database, design metadata for public network information retrieval, and index public network information. For the required risk factors, develop a product model framework, build an initial information retrieval logic, and record the product category, name, core factor, search word suffix, and other required elements. Classify each product library by different product categories and names, and continuously update and supplement the relevant search factors.

[0035] Based on the above logic, determine the index association information retrieval path and search target, form a "metadata" index, and expand gradually through product demand and category classification storage.

[0036] Taking property insurance as an example, it is classified as property insurance - non-vehicle insurance - property damage - enterprise property insurance. According to its risk factor elements, search for information such as fire (keywords "fire", "burning", "firefighting", …), extreme weather (keywords "heavy rain", "typhoon", "warning", …), etc.

[0037] Taking agricultural meteorological index insurance as an example, it is classified as property insurance - agricultural insurance - index - agricultural meteorological index insurance. According to its risk factor elements, search for information such as temperature (keywords "temperature", "high temperature", …), rain (keywords "rainfall", "precipitation", …), etc.

[0038] Step 1.2, intelligent integration of multiple search engines;

[0039] Compared with internal data docking platforms, public network data sources are diverse in structure. In order to take advantage of the data size and dimension, through the intelligent integration of multiple search engines, the difference retrieval logic of numerical, text, image, audio and video is realized to record the required data.

[0040] By providing an extensible adaptive data interface, accommodate multiple public network information sources to ensure the breadth of data collection as much as possible. At the same time, through interface constraints, confirm the range of time, type and other elements of the required information to improve the accuracy of data collection and ensure the efficiency of data application.

[0041] Tool integration: Through traditional data platforms, integrate various search engine APIs, deploy distributed crawlers, form a public information network data framework, and expand the coverage.

[0042] Condition association: associate information retrieval logic, based on product demand, determine search objects, categories, and other constraints, and develop a basic work framework.

[0043] Intelligent scheduling: combined with search keywords, set task scheduler, dynamically allocate real-time applications.

[0044] Step 2, public network information database design:

[0045] Build a 0-level database (original database) to store raw data collected from public networks. Build a 1-level database (calibration database) to calibrate the multi-dimensional availability of raw data in terms of quality level, integrity level, and information quantity level. Build a 2-level database (application database) to store insurance pricing available data after cleaning, fusion, and mining governance. Specifically, it includes:

[0046] Step 2.1, 0-level database (original database) construction;

[0047] For existing products, to combine internal data platforms, expand through service interfaces based on past platforms, form public data interfaces and differentiated groups, store and retrieve information to achieve application; for innovative products, the 0-level database is independent, and the application is realized through independent interfaces.

[0048] The public network data stored in the 0-level database includes structured data such as numerical values, as well as unstructured or semi-structured data such as text, images, audio, and video. It allows the aggregation of various risk factor data to build a comprehensive 0-level database (see Appendix Figure 3 ), or build a single-factor 0-level sub-database for each risk factor data, and the collection of all 0-level sub-databases as the 0-level database.

[0049] Step 2.2, data availability calibration rule design and 1-level database (calibration database) construction;

[0050] In existing pricing models, insurance companies rely heavily on third-party platform internal data, and can only submit data suggestions through communication, lacking data evaluation rules, and unable to evaluate data independently based on front-end business needs. Therefore, in the pricing process, internal data is usually assumed to be reliable, and a unified model is built, or manual adjustments are made based on subjective judgments.

[0051] Step 2.2.1, data availability calibration rule design;

[0052] Public network information sources are diverse in structure and availability, and must be calibrated and managed accordingly before being used for pricing models. Therefore, the invention designs data calibration rules for raw data obtained from public networks. Based on the preset rules, the information quality level, information integrity level, and information quantity level of the raw data are calibrated in multiple dimensions.

[0053] The data calibration rules are set as follows:

[0054] ① Information quality level: evaluate the reliability of data, set "reliable", "relatively reliable", "basically reliable", "unreliable" four levels. The basis is the reliability of data source. Including: 1) authority and transmission path of information source; 2) logical rationality of information content, completeness of evidence chain and identification of human manipulation; 3) transparency of information platform operation, content supervision and industry relevance, etc. Example: government website release > business agency release > scientific research agency release > self-media release.

[0055] ② Information integrity level: evaluate the matching level of data coverage to application requirements, set five, four, three, two, one five levels. Information integrity level is related to application requirements, and the specific calculation is: collected data coverage / application requirement data coverage rounded. For example: the model requires annual average temperature. If there are 11 months of average temperature, the integrity is 5; if there are 6 months of average temperature, the integrity is 3; if there are 2 months of average temperature, the integrity is 1.

[0056] ③ Information quantity level: evaluate the depth of application information provided by data. Set "quantitative", "semi-quantitative", "qualitative" three levels. For example: information is specific numerical value for quantitative, information is data range for semi-quantitative, information is qualitative description for qualitative.

[0057] Data labeling rules can be further adjusted and improved according to public network information and the expansion of application requirements.

[0058] Step 2.2.2, 1 level database (labeled database) construction;

[0059] According to the "data labeling rules", the usability of 0 level database data is judged from the above several dimensions (which can be further expanded), and each data and each dimension is marked and stored to form a 1 level database (refer to Figure 4 );

[0060] In implementation, on the basis of the original database, the pre-written labeling rules are automatically associated by the background service, and the information quality level, information integrity level and information quantity level of public network data are evaluated and labeled. Form a labeled database through automatic storage for subsequent processes.

[0061] Step 2.3, data governance method establishment and 2 level database (application database) construction;

[0062] Step 2.3.1, data governance method establishment;

[0063] Establish a public network information data governance method to make data available for insurance pricing modeling. Specifically as follows:

[0064] For unstructured and semi-structured data such as text, images, audio, and video, integrate intelligent tools such as NLP technology, computer vision, and speech recognition to extract the required information and convert it into structured data.

[0065] For semi-quantitative and qualitative level calibration data, aim for quantification through information mapping, machine learning, and deep learning to mine information and improve the information quantity of each group of data. The information quality level and information completeness level remain unchanged. For example, based on historical statistics of the relationship between water quality level and specific water quality index value, deduce from water quality level to water quality index value.

[0066] For incomplete information, aim for complete matching of application requirements by processing data (combination, extrapolation, and interpolation), AI comprehensive analysis (Kalman filter, neural network), and other methods to "data-level" fusion (not using decision-level fusion to avoid information loss) related incomplete data groups. The information quality level is the weighted average of the information quality level of each group of data according to the information completeness level.

[0067] Set the threshold conditions for the number of independent final information sources and the relative deviation of the data. For each group of calibration data that meets the set source number and deviation threshold conditions, and has the same information quality level (except for unreliable level), each group is upgraded by one information quality level.

[0068] Based on the above governance, data cleaning is performed to remove data with information quantity level of semi-quantitative and qualitative, completeness level less than 3, and quality level of unreliable. The remaining data is later marked with availability level and stored.

[0069] Step 2.3.2, 2-level database (application database) construction;

[0070] Apply data governance methods to 1-level database data for data format conversion, information quantity mining, information data-level fusion, information quality calibration upgrade, and data cleaning to improve data usability. Generate structured data with quality level above basic reliable and above, completeness level above 3, and information quantity level of quantitative with availability level marking, store, and form a 2-level database (e.g. Figure 5 ).

[0071] Step 3, collaborative pricing modeling:

[0072] Integrate public network information 2-level database (application database) data with different availability levels and insurance company internal information database data considered reliable to build a weight-type insurance product pricing model based on data availability level weighting. Specifically as shown in Figure 6 .

[0073] 1) Apply the database data of public network information to the weight assignment rule according to its final data availability level, the higher the level, the greater the value. Specifically:

[0074] The weight of the quality level "reliable" is 1, the weight of the "relatively reliable" is 0.8, and the weight of the "basically reliable" is 0.6.

[0075] The weight of the integrity level "5" is 1, the weight of the level "4" is 0.8, and the weight of the level "3" is 0.6.

[0076] After mining and cleaning, the "quantitative" level data is retained, and the weight is 1.

[0077] The data quality weight, the integrity weight and the information weight are multiplied to obtain the final sample data availability comprehensive weight. According to the feedback of the model application effect, it is determined whether the assignment rule needs to be adjusted.

[0078] 2) The public network application database is associated with the internal database to form an intermediate result data table combining internal and external data for model pricing application.

[0079] 3) The public network database and internal database data related to the target product are divided into training set and test set, the existing dimension factor is used as the model feature, and the modeling work is implemented based on the related platform to measure and estimate the loss risk.

[0080] 4) In the model measurement process, compared with the past sample (arithmetic mean) risk loss function based on internal data, the weight type risk loss function model suitable for public network data is established as follows:

[0081] ,

[0082] Wherein, is the predicted risk loss of a factor, is the sample loss value of the factor, is the weight of The weight of public network data is the comprehensive assignment of its availability level, and the weight of internal network sample is set to 1 when internal and external network collaborative pricing modeling is used.

[0083] If the predicted risk loss of internal data for a certain factor is , containing samples; at the same time, the external network adds samples for the factor, and the sample loss values are , and the availability comprehensive weight is , then the weight type risk loss function model of internal and external network collaboration is:

[0084] ,

[0085] wherein, is the predicted risk loss of the factor internal and external network cooperation.

[0086] The weight here specifically refers to a sample data multidimensional availability level comprehensive weight, which is used to measure factor pure risk estimation, and does not involve weight adjustment of the insurance company according to market and customer factors for the rate.

[0087] 5) According to the risk loss function of each factor involved in the insurance product, the total risk loss function of the product is calculated. The specific model form is determined by the insurance product design, which has linear type, nonlinear type, and covariance matrix type, etc. The present application can be directly applied.

[0088] 6) For innovative products lacking internal network data, only public network data interfaces are set, and the public network information application database directly performs modeling operations. For traditional business products, if the public network has abnormal security, etc., the public network data interface is cut off, and the pricing modeling based on the internal network data is returned. Regardless of the situation, it can be regarded as a special case of the weight type pricing model.

[0089] Taking agricultural meteorological index insurance as an example, the past internal database risk factors include crop air humidity and temperature. Through external public network information collection, environmental pollution factors and waterlogging factors can be included.

[0090] Step 4, model automatic optimization:

[0091] The traditional insurance product pricing only forms a fixed model result, and the modeling process is measured by artificial repetition every half year / year period, and new data is included in the old framework mode, and the factors are adjusted by artificial judgment. Therefore, in the traditional mode, the model iteration update speed is slow, which cannot meet the market rapid response demand, and at the same time, a large amount of human and time cost is consumed.

[0092] The present application to the established according to the data availability level weighting weight type insurance product pricing model, real-time monitoring front-end data source change and model output effect, by system background realizes the continuous iteration and closed loop optimization.

[0093] Specifically includes the following two types of ways (such as Figure 7 (a)-(b) shown in the figure, wherein (a) is an automatic monitoring data source change iteration update principle diagram, (b) is an automatic monitoring model application effect, backtracking optimization principle diagram):

[0094] First, automatically monitor the data source change, and iteratively update the model;

[0095] 1) The user sets the update frequency requirement (such as every day, every week), and the system background connects the public data interface to automatically refresh periodically. Once new data sources are found or existing data sources present data updates, step 1 is started to collect data.

[0096] 2) Steps 2 to 3 are automatically executed to sequentially update the public database data at all levels, continuously optimize the pricing model, and match the latest front-end data conditions.

[0097] Second, automatically monitor the application effect of the model and backtrack the optimization rules.

[0098] 1) The model results are written into a table and the backtracking monitoring logic is written, continuously monitoring and comparing the model prediction risks and actual risks to determine the difference characteristics.

[0099] 2) For cases where the difference exceeds the business warning limit, extract and attribute analysis is performed to locate the deviation characteristics and circle the sample data scope to be optimized.

[0100] 3) Optimize the data availability level assignment logic, simulate the preset model training and output the effect comparison, and confirm whether the model optimization results meet the business requirements.

[0101] 4) If not, backtrack and optimize the data governance link associated with the availability level improvement and cleaning logic, simulate the preset model training and output the effect comparison, and confirm whether the model optimization results meet the business requirements.

[0102] 5) If not, backtrack and optimize the data calibration link associated with the availability rule logic, simulate the preset model training and output the effect comparison.

[0103] 6) Determine the best pricing model under the existing data and factor dimension.

[0104] 7) According to actual business requirements, add new factor dimensions and expand the front-end data interface to refine the model risk assessment architecture.

[0105] In summary, the present application integrates public information platform data into the pricing modeling process, combines it with internal information platform data, expands the application field of insurance services, and improves the application capability of insurance services. A complete data framework, retrieval logic, data collection, data calibration, data governance, collaborative modeling, and model optimization seven-link full-process architecture is constructed. Compared with traditional insurance product pricing models, dynamic updating and optimization are achieved, and model accuracy is improved without the need for periodic manual calculations.

[0106] The public information application service provided by the application can continuously expand the data interface, and integrate new multi-source heterogeneous data sources into the model database. Compared with the fixed internal data docking platform of the traditional pricing model, it has the advantages of more extensive information and more timely updates. At the same time, the three-level database mode iteration of raw data, calibration data and application data is adopted to avoid data contamination, and the comprehensive application of public data and internal data can be realized for innovative products and traditional products.

[0107] Compared with the traditional product internal data platform for unified processing of standard format data, the public network data is widely sourced from professional business websites, scientific research websites to self-media, and the information structure is not the same, not only the internal network structured data form used in traditional insurance product pricing, but also unstructured and semi-structured data forms such as text, image, audio and video. The application stores multi-source heterogeneous data by integrating intelligent application tools and combining keyword filtering logic, uses differentiated scheduling methods to capture all the required information, and expands the data breadth.

[0108] Compared with internal network data, public network data has the characteristics of multi-source heterogeneity and mixed fish and dragon. It is necessary to evaluate and improve the data availability in order to make full use of it and fully exert its benefits. The application designs data availability grade division rules from multiple dimensions such as information quality, information integrity and information quantity, and judges the availability of the collected data according to the rules. Through the calibration of public network acquisition data, the product factor requirements can be directly locked according to the front-end business feedback to expand the retrieval and acquisition of data. On this basis, the data availability is improved through targeted governance methods in each dimension. Through the availability calibration and availability improvement process of public network acquisition data, the advantages of wide and large amount of public network information are fully utilized, and the accuracy of the pricing model is improved.

[0109] Compared with the traditional insurance product pricing model which is manually updated and optimized every six months or a year, the application realizes automatic detection of the system, and can realize real-time iterative optimization model combined with writing optimization logic. Thus, the pricing model can accurately reflect the latest market risk scenario, and improve its precision and application value.

[0110] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the application. It should be understood that the above description is only a specific embodiment of the application and is not intended to limit the application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application should be included in the protection scope of the application.

Claims

1. An intelligent pricing method for insurance based on public network information, characterized in that, The method comprises: Step 1, public network information data retrieval: according to the insurance product type and risk factor demand, a multi-source heterogeneous data retrieval logic is constructed, and original data is acquired through a public network; Step 2, three-level public network information database design: an original database is constructed for storing the original data; a calibration database is constructed for calibrating the usability of the original data, including: designing data calibration rules, calibrating the information quality level, information integrity level and information quantity level of the original data in multiple dimensions, labeling and storing to form a calibration database; an application database is constructed for data governance and storage of the calibration database data, including: For unstructured and semi-structured calibration data, the required information is intelligently extracted and converted into structured data; For calibration data with semi-quantitative and qualitative information quantity levels, quantitative information is intelligently mined to improve the information quantity level of each group of data; For calibration data with insufficient information integrity level, intelligent data fusion is performed on related incomplete data groups to form data that is completely matched with application requirements; Set the threshold conditions of the number of independent final information sources and the relative deviation of the data, and for each group of calibration data that meets the threshold conditions and has the same information quality level, each group is improved by one information quality level; Data cleaning is performed, the remaining data is re-marked and stored according to the usability level to form an application database; Step 3, cooperative pricing modeling: combine the public network information application database data with the insurance company internal database data to construct a weight type insurance product pricing model according to the data availability level; wherein the weight type insurance product pricing model adopts the following weight type risk loss function: assuming that the internal data predicts the risk loss of one factor as , containing samples; at the same time, the public network adds samples to the one factor, and the sample loss values are respectively , and the availability comprehensive weight is respectively , then the weight type risk loss function model of internal and external network cooperation is: , wherein, is the predicted risk loss for the one of the factors for the cooperation of the internal network and the external network; Step 4, model automatic optimization: real-time dynamic monitoring of public network data source changes and weight type insurance product pricing model application effect, automatic updating of the pricing model through the feedback mechanism, adjustment of data calibration and model weighting rules, realization of closed-loop optimization and continuous iteration of the pricing model. 2.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, An insurance product pricing modeling system combining public network information and internal network information is constructed, specifically including: A technical architecture for public network information insurance intelligent pricing application is constructed to provide scalable multi-source heterogeneous data support for insurance product pricing and dual-network services of public network information and internal network information that meet security requirements; Design a full-process technology for data collection, data calibration, data governance, data storage, data pricing application modeling and model optimization. 3.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, Step 1 includes: establishing a product classification system according to the type of each product, setting risk factors and retrieval keywords for each product, and forming a retrieval metadata index; based on the index, structured, unstructured and semi-structured data is collected from the public network through distributed crawlers and search engine APIs. 4.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, In step 2, the original database is constructed to store the original data collected from the public network, including: For traditional products, a public data interface and differentiated groups are formed through service interface expansion to store retrieval information for application; for innovative products, the original database is independently grouped and applied through an independent interface. 5.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, In step 2, the information quality level is used to evaluate the reliability of the data, the information integrity level is used to evaluate the matching level of the data coverage to the application requirements, and the information quantity level is used to evaluate the depth of the application information provided by the data: wherein, The information quality level is set as "reliable", "relatively reliable", "basically reliable" and "unreliable"; the information integrity level is set as 5, 4, 3, 2 and 1; and the information quantity level is set as "quantitative", "semi-quantitative" and "qualitative". 6.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, The step 3 comprises: The public information application database data is applied to establish a comprehensive weight assignment rule of availability in its final multidimensional data availability level, and the higher the level is, the greater the value is; The public information application database is associated with the internal database to form intermediate result data table of internal and external data combination for application of the pricing model; The public information application database data and the internal database data related to the pricing target product are divided into training set and test set to establish a collaborative pricing model to estimate loss risk. 7.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, When only public data is used, the weight type risk loss function model suitable for public data is: wherein, a predicted risk loss for one factor, y i a sample loss value for the one factor, i = 1, 2, …, N is the availability comprehensive weight of ; wherein, when modeling the internal and external network collaborative pricing, the internal network sample weight is set to 1.​ 8.The insurance intelligent pricing method based on public network information according to claim 1, characterized in that, The step 4 comprises: The public data source is automatically monitored, once new data source is found or the existing data source presents data update, the step 1-step 3 is started to sequentially update the public database at each level to realize continuous optimization of the pricing model and match the latest data condition of the front end; The application effect of the weight type insurance product pricing model is automatically monitored, and for the difference between the model prediction risk and the actual risk exceeding the business alert limit, the data availability level calibration, management and weight assignment are traced back to optimize to improve the application effect of the model.

Citation Information

Patent Citations

  • Intelligent insurance risk assessment and pricing system based on artificial intelligence

    CN116823496A

  • Intelligent campus operation and maintenance data management method and system based on big data

    CN118503910A

  • Intelligent premium pricing method and system based on multi-dimensional data analysis driving

    CN118887021A