Insurance fraud behavior detection and early warning system based on large language model

By applying an insurance fraud detection and early warning system based on a large language model in insurance companies, combined with multi-dimensional disaster risk assessment, the high fraud risk problem caused by the surge in claims after the disaster is solved, and efficient fraud detection and early warning are achieved.

CN120013682AInactive Publication Date: 2025-05-16STATE GRID ANHUI ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411820812.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The surge in claims after the disaster has caused insurance companies to face high risk of fraud, and it is difficult for existing technologies to effectively detect and early warning.

Method used

The insurance fraud detection and early warning system based on the large language model is adopted, and the application of the regional risk conditions is established module, the disaster risk prediction module, the disaster assessment module and the comprehensive assessment module, and the disaster risk index, the social vulnerability index and the emergency response capacity index are carried out to conduct multi-dimensional assessment and early warning.

Benefits of technology

It has achieved efficient detection and early warning of potential fraud in post-disaster claims, reduced the fraud risk of insurance companies, and improved the efficiency and accuracy of claims review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013682A_ABST
    Figure CN120013682A_ABST
Patent Text Reader

Abstract

The invention discloses an insurance fraud behavior detection and early warning system based on a large language model, and particularly relates to the field of insurance, and the system comprises a regional risk condition establishment module, a disaster risk pre-judgment module, a disaster evaluation module and a comprehensive evaluation module. High-risk regions are identified through parameters such as a disaster risk index, a social vulnerability index and an emergency response capability index, and early warning is carried out according to prediction data before a disaster occurs; after the disaster, the system can further screen out possible fraudulent behaviors by combining the insurance market permeability and fraudulent risk assessment, and intervention is carried out in time; the system intelligently judges and analyzes potential abnormal claim settlement requests based on historical claim settlement data, post-disaster emergency resource allocation, population density and other multi-dimensional factors, so that insurance fraud behaviors are effectively prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of insurance technology, and more specifically, to an insurance fraud behavior detection and early warning system based on a large language model. Background Art

[0002] Insurance fraud refers to the insured, beneficiary or third party defrauding the insurance company of compensation or benefits that should not be paid by means of false statements, forged evidence or concealment of facts. Insurance fraud is also a manifestation of insurance fraud, which usually refers to the insured falsely reporting losses, exaggerating damages or forging accidents during the claims process to obtain claims amounts that do not conform to the actual situation. Such behavior not only violates the principle of good faith, but also has a serious impact on the stability and fairness of the insurance industry.

[0003] After a disaster, the number of insurance claims usually increases dramatically. This is because the losses and impacts caused by the disaster cause a large number of policyholders to file claims with insurance companies. However, in the context of large-scale claims demand, insurance companies will face higher claims pressure, which also provides opportunities for insurance fraud and fraud. Due to the surge in the number of claims, insurance companies are more likely to have reduced processing efficiency and quality when reviewing and verifying each claim application, making it easier for some lawless elements to take advantage of the opportunity to commit fraud by means of forging evidence, falsely reporting losses or maliciously exaggerating losses. Therefore, the surge in claims after a disaster not only increases the operating risks of insurance companies, but also makes it more likely to trigger a large number of improper behaviors. Summary of the invention

[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an insurance fraud behavior detection and early warning system based on a large language model, which solves the problems raised in the above-mentioned background technology through the application of a regional risk condition establishment module, a disaster risk pre-determination module, a disaster assessment module and a comprehensive assessment module.

[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: an insurance fraud behavior detection and early warning system based on a large language model, comprising a regional risk condition establishment module, a disaster risk pre-determination module, a disaster assessment module and a comprehensive assessment module;

[0006] The regional risk condition establishment module is used to obtain regional information involved in the insurance business, and establish regional risk condition 1 and regional risk condition 2 for each region based on the regional information;

[0007] The disaster risk pre-determination module is used to determine each region in the regional information. If the regional risk condition 1 or the regional risk condition 2 is determined to be high risk, a disaster prediction model is established for the region;

[0008] The disaster assessment module includes a disaster prediction model, which is established by combining the disaster risk index, social vulnerability index and emergency response capability index; the disaster assessment value is calculated according to the disaster prediction model; if the disaster assessment value is higher than the disaster assessment threshold preset by the system, the area is marked as a first-level high-risk state, otherwise it is normal;

[0009] The comprehensive assessment module is used to establish a comprehensive assessment model for areas marked as level one high-risk status. The comprehensive assessment model is based on the disaster assessment value and is combined with the insurance market penetration assessment and fraud risk assessment. The comprehensive assessment value is obtained according to the comprehensive assessment model. If the comprehensive assessment value of the area is higher than the comprehensive assessment threshold preset by the system, the area will be marked as level two high-risk status.

[0010] In a preferred embodiment, the disaster risk index includes an assessment of the frequency of historical disasters and the rate of infrastructure damage; the social vulnerability index includes an assessment of population density; the emergency response capacity index includes an assessment of the density of emergency resource deployment before a disaster;

[0011] Regional risk condition 1: When the frequency of historical disasters is higher than the threshold of historical disasters preset by the system, and the population density is higher than the threshold of population density preset by the system, the area is judged to be high risk;

[0012] Regional risk condition two: When the infrastructure damage rate is higher than the infrastructure damage rate threshold preset by the system, and the pre-disaster emergency resource deployment density is higher than the pre-disaster emergency resource deployment density threshold preset by the system, the area is judged to be high risk.

[0013] In a preferred embodiment, the disaster prediction model included in the disaster assessment module is composed of a disaster risk index DRI, a social vulnerability index VSI, and an emergency response capability index ERI, and the disaster assessment value calculated by the disaster prediction model is set to D eval ;

[0014] DRI=αD f +βD i

[0015] VSI=γP d +δS i

[0016] ERI=E r +ζR t

[0017] D eval =λDRI+μVSI+νERI

[0018] Where D f is the frequency of historical disasters; D iis the infrastructure damage rate; α and β are weight coefficients, which are used to adjust D f and D i Contribution to the Disaster Risk Index (DRI);

[0019] Where P d is population density; S i is the social inequality index; γ and δ are weight coefficients, which are used to adjust P d and S i Contribution to the Social Vulnerability Index (VSI);

[0020] Where E r The density of emergency resources before disaster; R t is the historical post-disaster medical resource recovery time; and ζ are weight coefficients, which are used to adjust E r and R t Contribution to the Emergency Response Capability Index (ERI);

[0021] Among them, λ, μ, and ν are weight coefficients, which are used to adjust the contribution of disaster risk index DRI, social vulnerability index VSI, and emergency response capacity index ERI to the disaster prediction model;

[0022]

[0023] Where T eval Assess thresholds for disasters.

[0024] In a preferred embodiment, the comprehensive evaluation model includes an insurance market penetration evaluation model. pen , Fraud Risk Assessment risk ; Calculate the comprehensive evaluation value through the comprehensive evaluation model, and propose a comprehensive evaluation value of C eval ;

[0025] I pen =α′P ins +β′A subs

[0026] F risk =κC fraud +λ′T claims

[0027] C eval =ξD eval +μ′I pen +ν′F risk

[0028] Where P ins is the penetration rate of insurance products; subs is the consumer insurance participation; α ′ and β ′ are weight coefficients, which are used to adjust Pins and A subs Contribution

[0029] Among them C fraud is the proportion of historical fraud cases; T claims is the total amount of claims; κ and λ ′ are weight coefficients, which are used to adjust C fraud and T claims Fraud risk assessment risk The impact of

[0030] Among them, ξ, μ ′ , ν ′ are weight coefficients, which are used to adjust the disaster assessment value D eval 、Insurance Market Penetration Assessment I pen , Fraud Risk Assessment risk Contribution

[0031]

[0032] Where E eval Comprehensive evaluation threshold.

[0033] In a preferred embodiment, the historical disaster frequency D is calculated. f ;

[0034]

[0035] Among them, H i is the number of disaster events occurring in the i-th year; T i is the total potential number of disasters in year i; n is the total number of years calculated;

[0036] Calculate the infrastructure damage rate D i ;

[0037]

[0038] Where D ij is the damage degree of the jth infrastructure in the i-th disaster; I j is the total evaluation value of the jth infrastructure; m is the total number of infrastructures considered.

[0039] In a preferred embodiment, the population density P is calculated d ;

[0040]

[0041] Where P i is the population in the ith region; A j is the area of ​​the jth region; k is the total number of regions considered;

[0042] Calculate the social inequality index S i ;

[0043]

[0044] Where P k is the income or wealth level of the kth social group; P is the average income or wealth level of all groups in the region; l is the number of social groups.

[0045] In a preferred embodiment, the pre-disaster emergency resource deployment density E is calculated. r ;

[0046]

[0047] Where R i is the number of emergency resources of the i-th category; P i is the applicable population of the i-th type of resources; A is the area of ​​the region; n is the total number of emergency resource categories;

[0048] Calculate the historical post-disaster medical resource recovery time R t ;

[0049]

[0050] Where T i M is the time required for the recovery of medical resources after the i-th disaster; i is the target number of medical resources required for recovery after the i-th disaster; k is the number of post-disaster medical recovery events.

[0051] In a preferred embodiment, the insurance product penetration rate P is calculated ins ;

[0052]

[0053] in is the popularity of the i-th insurance product; M i is the total market population covered by the insurance product; n is the total number of insurance product types considered;

[0054] Calculating Consumer Insurance Participation A subs ;

[0055]

[0056] Where N subs,i is the number of insured persons of the i-th insurance product; P total is the total population of the region; m is the total number of insurance product types considered.

[0057] In a preferred embodiment, the historical fraud case ratio C is calculated.fraud ;

[0058]

[0059] where F i is the number of fraud cases in the i-th claim case; T j is the total number of j-th claim cases; n is the total number of claim cases considered;

[0060] Calculate the total amount of claims T claims ;

[0061]

[0062] Where Z i is the number of claims for the i-th category; V i is the single claim amount of the i-th type of claim request; m is the total number of claim request types.

[0063] Technical effects and advantages of the present invention:

[0064] 1. This solution is based on a multi-dimensional assessment of disaster risk, social vulnerability and emergency response capabilities, predicts high-risk areas in advance, and detects possible post-disaster claims fraud in combination with historical claims data, social factors and emergency resources. When large-scale claims demand occurs, the system can effectively identify abnormal claims behavior, reducing the fraud risk caused by the surge in post-disaster claims. Through comprehensive pre-disaster predictions and real-time monitoring, the efficiency of fraud detection in post-disaster claims can be improved, preventing insurance companies from falling into fraud traps due to the huge claims pressure brought by disasters.

[0065] 2. This solution combines disaster prediction models, insurance market penetration assessments and fraud risk analysis models, and provides comprehensive post-disaster claims predictions and fraud risk assessments through multi-dimensional data fusion. It breaks through the traditional model of a single parameter or a single perspective, making the assessment of high-risk areas more dynamic, especially after a disaster, it can promptly identify the areas with the most concentrated risks and promptly warn of possible insurance fraud.

[0066] 3. By establishing a comprehensive assessment model and combining the disaster assessment value with the insurance market penetration rate and fraud risk, the system can automatically adjust the risk assessment standards during the high-frequency claims stage after a disaster, and adjust the review process according to the specific risk conditions in each region. This function not only improves the efficiency of claims review, but also avoids the problem of insufficient manual review resources caused by the surge in claims after a disaster, allowing the system to respond to emergencies more efficiently and reduce delays and omissions in manual operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0068] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0069] Refer to the instruction manual Figure 1 , an insurance fraud behavior detection and early warning system based on a large language model according to an embodiment of the present invention comprises a regional risk condition establishment module, a disaster risk pre-determination module, a disaster assessment module and a comprehensive assessment module;

[0070] The regional risk condition establishment module is used to obtain regional information involved in the insurance business, and establish regional risk condition 1 and regional risk condition 2 for each region based on the regional information;

[0071] The disaster risk pre-determination module is used to determine each region in the regional information. If the regional risk condition 1 or the regional risk condition 2 is determined to be high risk, a disaster prediction model is established for the region;

[0072] The disaster assessment module includes a disaster prediction model, which is established by combining the disaster risk index, social vulnerability index and emergency response capability index; the disaster assessment value is calculated according to the disaster prediction model; if the disaster assessment value is higher than the disaster assessment threshold preset by the system, the area is marked as a first-level high-risk state, otherwise it is normal;

[0073] The comprehensive assessment module is used to establish a comprehensive assessment model for areas marked as a first-level high-risk state. The comprehensive assessment model is based on the disaster assessment value and is combined with the insurance market penetration assessment and fraud risk assessment. The comprehensive assessment value is obtained according to the comprehensive assessment model. If the comprehensive assessment value of the area is higher than the comprehensive assessment threshold preset by the system, the area will be marked as a second-level high-risk state;

[0074] The above scheme is based on a hierarchical multi-dimensional disaster and insurance risk assessment model. By combining disaster prediction and insurance fraud detection, it ensures that high-risk areas can be warned and monitored in real time and dynamically, especially potential fraud in the post-disaster claims process. In this scheme, firstly, by setting regional risk conditions, high-risk areas that need attention are screened out. This step is based on a comprehensive judgment based on multi-dimensional indicators such as historical disaster frequency, infrastructure vulnerability, and population density to ensure that the early prediction of disaster risks is more accurate. Next, a disaster prediction model is established for these high-risk areas. It combines the disaster risk index, social vulnerability index, and emergency response capability index. By calculating the disaster assessment value, the area is divided into different risk levels, specifically to the first-level high-risk state and normal state. This step is to warn high-risk areas before the disaster occurs. For areas that have been marked as high-risk, by further establishing a comprehensive assessment model, introducing the penetration rate of the insurance market and fraud risk assessment, weighting is performed on the basis of the disaster assessment value to form a comprehensive assessment value, so as to determine whether the area needs to be further monitored.

[0075] In addition, the working principle of this solution is to conduct comprehensive monitoring of high-risk areas that may trigger large-scale claims in advance through pre-disaster risk assessment, and to more effectively identify potential fraud after the disaster; through this forward-looking assessment and early warning mechanism, not only the efficiency of claims is optimized, but also the probability of insurance companies encountering fraud risks when making large-scale post-disaster claims is effectively reduced, and the operating efficiency of the entire system and the ability to respond to disasters are improved; the core advantage of this design is to combine disaster risks with the early warning system of insurance fraud detection, and to improve accuracy through multi-level and multi-dimensional assessments, avoiding the limitations of a single perspective or static analysis, thereby achieving more accurate risk control and management; the advantage of this solution is not only reflected in improving the efficiency of post-disaster response, but also enabling insurance companies to formulate corresponding claims and prevention strategies before a disaster occurs, thereby saving the company a lot of potential costs and risks;

[0076] In this system, the large language model can be applied to multiple links such as post-disaster claims analysis, insurance fraud detection, data processing and early warning. Specifically, the application of the large language model is mainly reflected in the following aspects:

[0077] Claims data analysis and document processing: The big language model can process massive amounts of claims data and documents, such as claims applications, accident reports, and customer statements. Through natural language processing (NLP) technology, the model can extract key information from documents, such as accident type, claim amount, and customer history, and identify potential anomalies or contradictions. For example, whether there are inconsistencies in the accident details described by the customer during the claims process, or whether there are duplicate claims applications, all of which can be quickly discovered through the text understanding capabilities of the big language model.

[0078] Fraud detection: When detecting insurance fraud, the big language model can learn the language characteristics of normal and fraudulent claims by analyzing a large number of historical claims records. For example, fraud cases may involve false descriptions, forged evidence, or exaggerated loss amounts. The big language model can detect these abnormal patterns in real time through text comparison and semantic analysis, and issue timely warnings. This approach is more flexible and accurate than traditional rule-based fraud detection systems.

[0079] Predicting post-disaster claims demand: The big language model can also be used to predict post-disaster claims demand. After a disaster occurs, the system can analyze relevant news reports, social media content, government-issued post-disaster reports and other text information, combined with historical data, to help predict which regions will face higher claims requests after the disaster. This text-based big data analysis capability can provide decision support for post-disaster early warning and resource allocation.

[0080] Cross-domain data fusion: The large language model can also combine data from different sources, such as meteorological data, social media information, historical disaster records, customer information, etc., and comprehensively assess disaster risk and fraud risk through multimodal learning (including text and structured data). For example, the model can automatically understand and integrate text information output by the disaster prediction model (such as climate change reports) and insurance claims data, thereby providing the system with a more comprehensive assessment of high-risk status.

[0081] In short, the application of the big language model in this system is to automatically process and analyze claims documents, detect fraud, predict post-disaster claims needs, and provide multi-dimensional data support through its powerful natural language understanding and text generation capabilities; this not only improves the intelligence level of the system, but also makes post-disaster response more efficient and accurate;

[0082] The large language model in the above solution can be an open source large language model, including but not limited to GPT-3 (accessed through OpenAI's API) or BERT (launched by Google, suitable for text understanding and classification tasks) or RoBERTa (an improved version of BERT). These models can be used to process claims documents, fraud detection, and post-disaster information analysis. The specific choice depends on task requirements and computing resources.

[0083] The disaster risk index includes the assessment of the frequency of historical disasters and the rate of infrastructure damage; the social vulnerability index includes the assessment of population density; the emergency response capacity index includes the assessment of the density of emergency resource deployment before the disaster;

[0084] Regional risk condition 1: When the frequency of historical disasters is higher than the threshold of historical disasters preset by the system, and the population density is higher than the threshold of population density preset by the system, the area is judged to be high risk;

[0085] Regional risk condition two: When the infrastructure damage rate is higher than the infrastructure damage rate threshold preset by the system, and the pre-disaster emergency resource deployment density is higher than the pre-disaster emergency resource deployment density threshold preset by the system, the area is judged to be high risk.

[0086] The disaster prediction model included in the disaster assessment module is composed of the disaster risk index DRI, the social vulnerability index VSI, and the emergency response capability index ERI, and the disaster assessment value calculated by the disaster prediction model is proposed to be D eval ;

[0087] DRI=αD f +βD i

[0088] VSI=γP d +δS i

[0089] ERI=E r +ζR t

[0090] D eval =λDRI+μVSI+νERI

[0091] Where D f is the historical disaster frequency (unit: number of disasters per year), reflecting the frequency of disasters in a certain area; D i is the infrastructure damage rate (unit: percentage), which is used to measure the extent of infrastructure damage; α and β are weight coefficients, which are used to adjust D f and D i Contribution to the Disaster Risk Index (DRI);

[0092] Where P d S is the population density (unit: person / km2), which reflects the number of residents per km2 in the region. The higher the population density, the more difficult it is to recover from a disaster. i is the social inequality index (unit: index value), which is used to measure the degree of social inequality in a region. The higher the value, the more difficult it is to recover from a disaster. γ and δ are weight coefficients, which are used to adjust P d and S i Contribution to the Social Vulnerability Index (VSI);

[0093] Where E r R is the density of emergency resource deployment before a disaster (unit: number of emergency resources per thousand people), which reflects the pre-disaster resource preparation of the region. The higher the resource density, the more effective the disaster response. tis the historical post-disaster medical resource recovery time (unit: hour), which reflects the speed of post-disaster medical resource recovery in the region in historical data. The faster the recovery, the smaller the post-disaster impact; and ζ are weight coefficients, which are used to adjust E r and R t Contribution to the Emergency Response Capability Index (ERI);

[0094] Among them, λ, μ, and ν are weight coefficients, which are used to adjust the contribution of disaster risk index DRI, social vulnerability index VSI, and emergency response capacity index ERI to the disaster prediction model;

[0095]

[0096] Where T eval It is the disaster assessment threshold, which is a standard value used to determine whether a region is in a high-risk state.

[0097] Comprehensive evaluation model includes insurance market penetration assessment I pen , Fraud Risk Assessment risk ; Calculate the comprehensive evaluation value through the comprehensive evaluation model, and propose a comprehensive evaluation value of C eval ;

[0098] I pen =α′P ins +β′A subs

[0099] F risk =κC fraud +λ′T claims

[0100] C eval =ξD eval +μ′I pen +ν′F risk

[0101] Where P ins is the insurance product penetration rate (unit: percentage), which is used to reflect the popularity of insurance products in the region; A subs is the consumer insurance participation rate (unit: insurance rate), which is used to measure the willingness of regional consumers to purchase insurance when facing disasters; α ′ and β ′ are weight coefficients, which are used to adjust P ins and A subs Contribution

[0102] Among them C fraud is the historical fraud case ratio (unit: case ratio), and the fraud case ratio in a certain region is calculated based on historical data; T claimsis the total amount of claims (unit: total amount of claims), which is used to reflect the amount of claims in the region. Too much claims may be accompanied by fraud risks; κ and λ ′ are weight coefficients, which are used to adjust C fraud and T claims Fraud risk assessment risk The impact of

[0103] Among them, ξ, μ ′ , ν ′ are weight coefficients, which are used to adjust the disaster assessment value D eval 、Insurance Market Penetration Assessment I pen , Fraud Risk Assessment risk Contribution

[0104]

[0105] Where E eval The comprehensive assessment threshold is a standard value used to determine whether a region is in a level 2 high-risk state.

[0106] Calculate the frequency of historical disasters D f ;

[0107]

[0108] Among them, H i is the number of disaster events occurring in the i-th year; T i is the total potential number of disasters in the i-th year, such as the number of possible events such as extreme weather and earthquakes; n is the total number of years calculated; the frequency of historical disasters D f By calculating the frequency of disasters each year and standardizing the frequency of disasters each year to a scale of 100, the frequency of disasters in the region can be calculated. The ratio of the number of disasters each year to the number of potential disasters that year represents the disaster occurrence rate, which is then weighted for use in the total cycle.

[0109] Calculate the infrastructure damage rate D i ;

[0110]

[0111] Where D ij is the damage degree of the jth infrastructure in the i-th disaster (unit: damage degree percentage); I j is the total assessment value of the jth infrastructure (unit: original value of infrastructure or estimated value); m is the total number of infrastructures considered; the above formula calculates the weighted average proportion of all infrastructures damaged in multiple disasters. The degree of damage to each infrastructure is proportional to its original assessment value. The overall infrastructure damage rate is obtained by weighting the damage proportions of all infrastructures.

[0112] Calculate the population density P d ;

[0113]

[0114] Where P i is the population in the ith region; A j is the area of ​​the jth region; k is the total number of regions considered; the above formula calculates the population density of the overall region by considering the population and area of ​​multiple small regions. The population density of each region is calculated by the ratio of its total population and area, and then the population density of the overall region is obtained by summing up the density values ​​of each region;

[0115] Calculate the social inequality index S i ;

[0116]

[0117] Where P k is the income or wealth level of the kth social group; P is the average income or wealth level of all groups in the region; l is the number of social groups; the above formula measures the deviation of the wealth or income of each social group from the average level, and expresses the degree of social inequality. By calculating the square of the deviation of each group from the average level and then weighting it to all groups, the overall social inequality index is finally obtained.

[0118] Calculate the density of emergency resources before the disaster E r ;

[0119]

[0120] Where R i is the number of emergency resources of the i-th category; the density of emergency resources before the disaster is E r P in the formula i is the applicable population of the i-th type of resources; A is the area of ​​the region; the density of emergency resource deployment before the disaster is E r The n in the formula is the total number of emergency resource categories. The above formula is used to calculate the density of emergency resources in the region, taking into account the number of each type of emergency resources, the applicable population, and the distribution area. By weighting the population of each type of resource, it reflects the layout of emergency resources before the disaster.

[0121] Calculate the historical post-disaster medical resource recovery time R t ;

[0122]

[0123] Where T i M is the time required for the recovery of medical resources after the i-th disaster; iis the target number of medical resources required for recovery after the i-th disaster; the historical recovery time of medical resources after disasters R t The k in the formula is the number of post-disaster medical recovery events; this formula measures the time required for post-disaster medical resource recovery and takes into account the target number of resources required for recovery. The ratio of recovery time to the number of resources reflects the efficiency of recovery.

[0124] Calculate the insurance product penetration rate P ins ;

[0125]

[0126] in is the popularity of the ith insurance product (unit: number of insurance products); insurance product popularity rate P ins M in the formula i The total market population covered by the insurance product; the insurance product penetration rate P ins Where n is the total number of insurance product types considered; the formula is weighted by the market coverage of each insurance product to calculate the overall insurance product penetration. The penetration of each insurance product is proportional to the total population of the market it covers, and the penetration rate is comprehensively evaluated;

[0127] Calculating Consumer Insurance Participation A subs ;

[0128]

[0129] Where N subs,i is the number of insured persons of the i-th insurance product; P total is the total population of the region; consumer insurance participation rate A subs The m in the formula is the total number of insurance product types considered; the above formula calculates the ratio of the number of insured people of each insurance product to the total population of the region, comprehensively reflecting the overall insurance participation in the region.

[0130] Calculate the historical fraud case ratio C fraud ;

[0131]

[0132] where F i is the number of fraud cases in the i-th claim case; T j is the total number of claims cases of the jth period; the proportion of historical fraud cases C fraud Where n is the total number of claims considered; the above formula measures the proportion of fraud cases in the total claims by calculating the proportion of fraud cases in each claim case. The higher the value, the more frequent the fraud.

[0133] Calculate the total amount of claims T claims ;

[0134]

[0135] Where Z i is the number of claims for the i-th category; V i is the single claim amount of the i-th type of claim; the total number of claim requests T claims Where m is the total number of claim types. The above formula calculates the total amount of claims by weighting the amount of each type of claim. The total amount of claims is affected by the number and amount of each type of claim.

[0136] It should be noted that the above parameters are of great significance in the insurance fraud detection and early warning system, and can provide the system with multi-dimensional risk assessment and decision-making support. First, the frequency of historical disasters and the rate of infrastructure damage reveal the disaster risk of the region, help predict the surge in post-disaster claims demand, and then assess potential fraud risks. Population density and social inequality index reflect the vulnerability of the region. Regions with dense populations and high social inequality often face greater claims pressure, which is prone to excessive claims or fraud. The density of emergency resource deployment before the disaster and the historical post-disaster medical resource recovery time measure the ability of post-disaster resource response. Regions with poor recovery capabilities may lead to delays in claims and additional claims requests, increasing the occurrence of fraud. The penetration rate of insurance products and consumer insurance participation help assess the penetration of the insurance market and customers' insurance awareness. Regions with low penetration and participation may lead to false claims due to lack of effective insurance knowledge. The proportion of historical fraud cases directly measures the frequency of fraud in the region, while the total number of claims requests reflects the scale of regional claims activities. Both are important bases for judging fraud risks. Combining these parameters, the system can more accurately identify high-risk areas and predict abnormal behaviors in post-disaster claims, thereby effectively preventing and identifying insurance fraud.

[0137] This solution uses a multi-level disaster prediction and risk assessment mechanism, combined with specific parameters such as the disaster risk index, social vulnerability index, and emergency response capability index. It first identifies high-risk areas and predicts the surge in post-disaster claims demand in advance, thereby providing a basis for subsequent fraud detection. Through a comprehensive assessment model, the disaster assessment value is combined with the insurance market penetration rate and fraud risk assessment to monitor possible abnormal claims behavior in real time. The system uses historical claims data, the proportion of fraud cases, and the total number of claims requests to identify potential fraudulent behavior, and uses technologies such as large language models to intelligently analyze and detect documents to promptly warn of possible fraud risks. This process realizes a closed loop from early warning to real-time detection, helping insurance companies to effectively identify and intervene in high-risk behaviors before disasters occur.

[0138] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An insurance fraud detection and early warning system based on a large language model, comprising a regional risk condition establishment module, a disaster risk pre-determination module, a disaster assessment module and a comprehensive assessment module, characterized in that: The regional risk condition establishment module is used to obtain regional information involved in the insurance business, and establish regional risk condition 1 and regional risk condition 2 for each region based on the regional information; The disaster risk pre-determination module is used to determine each region in the regional information. If the regional risk condition 1 or the regional risk condition 2 is determined to be high risk, a disaster prediction model is established for the region; The disaster assessment module includes a disaster prediction model, which is established by combining the disaster risk index, social vulnerability index and emergency response capability index; the disaster assessment value is calculated according to the disaster prediction model; if the disaster assessment value is higher than the disaster assessment threshold preset by the system, the area is marked as a first-level high-risk state, otherwise it is normal; The comprehensive assessment module is used to establish a comprehensive assessment model for areas marked as level one high-risk status. The comprehensive assessment model is based on the disaster assessment value and is combined with the insurance market penetration assessment and fraud risk assessment. The comprehensive assessment value is obtained according to the comprehensive assessment model. If the comprehensive assessment value of the area is higher than the comprehensive assessment threshold preset by the system, the area will be marked as level two high-risk status.

2. The insurance fraud detection and early warning system based on a large language model according to claim 1, characterized in that: The disaster risk index includes the assessment of the frequency of historical disasters and the rate of infrastructure damage; the social vulnerability index includes the assessment of population density; the emergency response capacity index includes the assessment of the density of emergency resource deployment before the disaster; Regional risk condition 1: When the frequency of historical disasters is higher than the threshold of historical disasters preset by the system, and the population density is higher than the threshold of population density preset by the system, the area is judged to be high risk; Regional risk condition two: When the infrastructure damage rate is higher than the infrastructure damage rate threshold preset by the system, and the pre-disaster emergency resource deployment density is higher than the pre-disaster emergency resource deployment density threshold preset by the system, the area is judged to be high risk.

3. The insurance fraud detection and early warning system based on a large language model according to claim 2 is characterized by: The disaster prediction model included in the disaster assessment module is composed of the disaster risk index DRI, the social vulnerability index VSI, and the emergency response capability index ERI, and the disaster assessment value calculated by the disaster prediction model is proposed to be D eval ; DRI=αD f +βD i VSI=γP d +δS i YOU WERE= AND r +ζR t D eval =λDRI+μVSI+νERI Where D f is the frequency of historical disasters; D i is the infrastructure damage rate; α and β are weight coefficients, which are used to adjust D f and D i Contribution to the Disaster Risk Index (DRI); Where P d is population density; S i is the social inequality index; γ and δ are weight coefficients, which are used to adjust P d and S i Contribution to the Social Vulnerability Index (VSI); Where E r The density of emergency resources before disaster; R t is the historical post-disaster medical resource recovery time; and ζ are weight coefficients, which are used to adjust E r and R t Contribution to the Emergency Response Capability Index (ERI); Among them, λ, μ, and ν are weight coefficients, which are used to adjust the contribution of disaster risk index DRI, social vulnerability index VSI, and emergency response capacity index ERI to the disaster prediction model; Where T eval Assess thresholds for disasters.

4. The insurance fraud detection and early warning system based on a large language model according to claim 3 is characterized by: Comprehensive evaluation model includes insurance market penetration assessment I pen , Fraud Risk Assessment risk ; Calculate the comprehensive evaluation value through the comprehensive evaluation model, and propose a comprehensive evaluation value of C eval ; I pen =α′P ins +β′A subs F risk =κC fraud +λ′T claims C eval =ξD eval +μ′I pen +ν′F risk Where P ins is the penetration rate of insurance products; A subs is the consumer insurance participation; α ′ and β ′ are weight coefficients, which are used to adjust P ins and A subs Contribution Among them C fraud is the proportion of historical fraud cases; T claims is the total amount of claims; κ and λ ′ are weight coefficients, which are used to adjust C fraud and T claims Fraud risk assessment risk The impact of Among them, ξ, μ ′ , ν ′ are weight coefficients, which are used to adjust the disaster assessment value D eval 、Insurance Market Penetration Assessment I pen , Fraud Risk Assessment risk Contribution Where E eval Comprehensive evaluation threshold.

5. The insurance fraud detection and early warning system based on a large language model according to claim 4 is characterized by: Calculate the frequency of historical disasters D f ; Among them, H i is the number of disaster events occurring in the i-th year; T i is the total potential number of disasters in year i; n is the total number of years calculated; Calculate the infrastructure damage rate D i ; Where D ij is the damage degree of the jth infrastructure in the i-th disaster; I j is the total evaluation value of the jth infrastructure; m is the total number of infrastructures considered.

6. The insurance fraud detection and early warning system based on a large language model according to claim 5, characterized in that: Calculate the population density P d ; Where P i is the population in the ith region; A j is the area of ​​the jth region; k is the total number of regions considered; Calculate the social inequality index S i ; Where P k is the income or wealth level of the kth social group; is the average income or wealth level of all groups in the region; l is the number of social groups.

7. The insurance fraud detection and early warning system based on a large language model according to claim 6 is characterized by: Calculate the density of emergency resources before the disaster E r ; Where R i is the number of emergency resources of the i-th category; P i is the applicable population of the i-th type of resources; A is the area of ​​the region; n is the total number of emergency resource categories; Calculate the historical post-disaster medical resource recovery time R t ; Where T i M is the time required for the recovery of medical resources after the i-th disaster; i is the target number of medical resources required for recovery after the i-th disaster; k is the number of post-disaster medical recovery events.

8. The insurance fraud detection and early warning system based on a large language model according to claim 7, characterized in that: Calculate the insurance product penetration rate P ins ; Among them I pi is the popularity of the i-th insurance product; M i is the total market population covered by the insurance product; n is the total number of insurance product types considered; Calculating Consumer Insurance Participation A subs ; Where N subs,i is the number of insured persons of the i-th insurance product; P total is the total population of the region; m is the total number of insurance product types considered.

9. The insurance fraud detection and early warning system based on a large language model according to claim 8, characterized in that: Calculate the historical fraud case ratio C fraud ; where F i is the number of fraud cases in the i-th claim case; T j is the total number of j-th claim cases; n is the total number of claim cases considered; Calculate the total amount of claims T claims ; Where Z i is the number of claims for the i-th category; V i is the single claim amount of the i-th type of claim request; m is the total number of claim request types.