Insurance data auditing system based on large language model
Through the insurance data review system based on the large language model, multi-level information division, marking and abnormal detection of insurance claims data is solved, which solves the problem of difficult to deal with complex text data and multi-dimensional analysis in the existing technology, and realizes efficient and accurate insurance data review and risk management.
Patent Information
- Application Number
- CN202411820816.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing insurance data review technology is difficult to effectively process complex text data and analyze the authenticity of claims applications from multiple dimensions, resulting in misjudgment or misjudgment, making it difficult to meet the audit needs of complex or fraud risks.
The insurance claim data review system based on a large language model is adopted, and the insurance claim data is structured and multi-dimensionally analyzed through modules such as data preprocessing, information layer division and marking, abnormal detection, risk score and early warning, feedback and model optimization.
It improves the accuracy of identifying abnormal data, reduces the rate of misjudgment, improves the accuracy and efficiency of insurance data review, and realizes multi-dimensional accurate audit and real-time risk monitoring.
Smart Images

Figure CN120013683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of insurance data review, and more specifically, to an insurance data review system based on a large language model. Background Art
[0002] Insurance refers to financial services that provide risk protection. Through insurance contracts, insurance companies promise to policyholders to pay compensation when an agreed risk event occurs. Insurance services cover a wide range of areas, including multiple fields. Its core purpose is to provide financial compensation to the insured or its stakeholders and reduce the financial losses caused by risk events. Insurance data refers to all information generated and accumulated in the insurance business process, covering data from policy signing, premium payment, claim application to claim settlement. Insurance data plays an important role in claim review, risk assessment and decision-making, and is the basis for insurance companies to prevent fraud, control risks and optimize operations.
[0003] In existing technologies, insurance data review often relies on manual experience and is unable to process complex text data in a targeted manner or analyze the authenticity of claims applications from multiple dimensions. This traditional review method has obvious shortcomings when dealing with complex or fraudulent risks, is prone to misjudgments or omissions, and is difficult to meet review needs. Summary of the invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an insurance data review system based on a large language model, which performs hierarchical analysis and recognition on insurance data through a large language model to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above-mentioned object, the present invention provides the following technical solutions: an insurance data review system based on a large language model, comprising a data preprocessing module, an information layer division and marking module, an anomaly detection module, a risk scoring and early warning module, and a feedback and model optimization module;
[0006] The data preprocessing module is used to obtain the text data of the claim application before the insurance claim is settled, and to clean and standardize the text data, delete redundant information, and obtain preliminary processed data;
[0007] The information layer division and labeling module divides and labels the preliminary processed data according to the information layer through the large language model. The information layer includes the identity feature information layer, the event description information layer, the loss description information layer, and the medical report or auxiliary certificate information layer;
[0008] The anomaly detection module compares the current application data with the historical application data through a model. If the identity feature information layer detection or the event description information layer detection is determined to be abnormal, the risk assessment is performed through the risk scoring and early warning module;
[0009] The risk scoring and early warning module generates risk scores based on analysis results, marks high-risk applications and triggers early warnings, and is used for manual review;
[0010] The feedback and model optimization module feeds back the results of manual review to the model, optimizes the model recognition rules, and is used to improve audit accuracy.
[0011] In a preferred embodiment, before the insurance claim is settled, the data preprocessing module obtains the text data of the claim application, cleans and standardizes the text data, deletes redundant information, obtains preliminary processed data, and proposes a preliminary processed data set D;
[0012] D={T i ∣T i =Clean(R i )andR i ∈Raw Data}
[0013] Where D represents the cleaned and standardized dataset; T i represents the i-th cleaned record; R i represents the ith raw data; Raw Data represents the raw data; the function Clean(·) is used to delete redundant information, standardize the text format, and standardize the data; the result of data preprocessing is a structured data set, which is used by the large language model of the information layer division and labeling module for classification and anomaly detection;
[0014] The information layer division and labeling module classifies the preliminary processed data set D through a large language model, and divides and labels the data into four information layers; the proposed identity feature information layer is L s , the event description information layer is L e , the loss description information layer is L d , medical report or supporting certificate information layer is L m ;
[0015] D={L s ,L e ,L d ,L m}
[0016] Where L s , L e , L d , L m Represents four information layers respectively.
[0017] In a preferred embodiment, the judgment condition of the identity feature information layer detection includes: when the identity information consistency parameter is lower than the identity information consistency parameter threshold preset by the system and the identity authentication frequency parameter is higher than the identity authentication frequency parameter threshold preset by the system, it is judged as abnormal;
[0018] The judgment conditions for event description information layer detection include: when the event description similarity parameter is higher than the event description similarity parameter threshold preset by the system and the event time consistency parameter is lower than the event time consistency parameter threshold preset by the system, it is judged as abnormal.
[0019] In a preferred embodiment, the identity feature information layer L s Including evaluating the identity information consistency parameter P si , authentication frequency parameter P sf ;
[0020] Identity information consistency parameter P si Include name matching sub-parameter M n , ID number matching sub-parameter M c , Address matching sub-parameter M a ;
[0021] P si Used to measure the consistency of the identity information of the applicant's name, ID number and address submitted with the historical records, so as to identify whether there is a risk of inconsistency in the identity information;
[0022] P si =α·M n +β·M c +γ·M a
[0023] Where P si It is used to measure the overall matching degree of identity information; α, β, and γ are the consistency weights of name, ID number, and address, respectively, which are used to balance the impact of different information on the overall consistency;
[0024]
[0025]
[0026]
[0027] Where N is the total number of name matches in the history; Indicates the name weight of the i-th record; Indicates whether the name of the i-th record matches; C is the total number of ID number matches in the historical records; Indicates the certificate number weight of the i-th record; Indicates whether the certificate number matches, 1 for a match and 0 for a mismatch; A is the total number of address matches in the history record; The weight of the address record; A flag indicating whether the address matches;
[0028] Authentication frequency parameter P sf Including contact information change frequency F p , Address change frequency F a ;
[0029]
[0030] Where P sf Used to assess the frequency of changes in contact information and address; c and σ a are the weight coefficients of contact information and address, respectively, to adjust their influence on the change frequency; T is the length of the observation time window;
[0031]
[0032]
[0033] Where P represents the number of contact records; Indicates whether the i-th record has changed; T p represents the total number of contact method changes during the observation period; A represents the number of address records; Indicates address change; T a Indicates the total number of address changes during the observation period.
[0034] In a preferred embodiment, the event description information layer L e Including evaluating the event description similarity parameter P es , event time consistency parameter P et ;
[0035] Event description similarity parameter P es Used to compare the similarity between the current event description and the historical record; event description similarity parameter P es Including keyword similarity S k , sentence structure similarity S t ;
[0036] P es =λ·S k +μ·S t
[0037]
[0038]
[0039] Among them, λ and μ are adjustment coefficients, respectively controlling the keyword similarity S k Similarity with sentence structure S t P es The influence weight of K is the total number of keywords; w is j is the weight of the jth keyword; is the matching tag of the jth keyword; M is the total number of sentences; is the matching tag of the mth sentence;
[0040] Event time consistency parameter P et Used to measure the reasonableness between the time when the event occurred and the time when the application was submitted;
[0041]
[0042] Where T s is the actual time of the event; T e T is the application submission time, that is, the time when the applicant submits the claim application; max is the maximum allowed time deviation.
[0043] In a preferred embodiment, the loss description information layer L d Including the floating parameter P of the estimated loss amount di , loss detail consistency parameter P da ;
[0044] Loss amount floating parameter P di Used to measure the difference between the loss amount of the current application and the historical record, and detect whether there is abnormal fluctuation in the amount; the loss detail consistency parameter P da Used to measure the consistency between current and historical loss details and detect whether there are differences in the detailed descriptions;
[0045]
[0046]
[0047] Among them, M current is the loss amount of the current claim; M average is the average loss amount of similar historical events; w i is the weight of the i-th loss detail; Indicates the matching of current and historical loss details; D is the total number of fields in the loss details;
[0048] Medical report or supporting evidence information layer L m Includes evaluation document source verification parameter P mr , file structure consistency parameter P mf ;
[0049] File origin verification parameter P mr Used to detect the source consistency of the attached certification documents; file structure consistency parameter P mf Used to measure the similarity of file formats and structures;
[0050]
[0051]
[0052] Where R is the total number of file source fields; w i is the weight of the i-th source field, which is used to adjust the impact of each source field on the overall source consistency; is the matching status of the i-th source field;
[0053] Where F is the total number of file structure fields; w j is the weight of the jth structure field; It is the matching status of the jth structure field.
[0054] In a preferred embodiment, the identity feature information layer risk score is set to R s ; Identity feature information layer risk score R s Based on the identity feature information layer L s The identity information consistency parameter P in si and the authentication frequency parameter P sf to measure risk;
[0055] R s =a·P si +b·P sf
[0056] Where R s It is used to reflect whether there are abnormalities in the consistency and update frequency of the applicant's identity information; a and b are adjustment coefficients that determine P si and P sf R s The weight of influence;
[0057] The risk score of the proposed event description information layer is R e ; Event description information layer risk score R e Based on the event description information layer L e The event description similarity parameter P in es and the event time consistency parameter P et to measure risk;
[0058] R e =c·P es +d·P et
[0059] Where R e It is used to reflect the repeatability and time consistency of event description; c and d are adjustment coefficients;
[0060] The proposed loss description information layer risk score is R d ; Risk score R of loss description information layer d Based on the loss description information layer L d The floating parameter P of the loss amount in di and loss detail consistency parameter P da To assess the consistency of loss amounts and detailed descriptions;
[0061] R d =e·P di +f·P da
[0062] Where R d It is used to reflect the fluctuation of loss amount and the consistency of loss details description; e and f are adjustment coefficients, which determine the impact of loss amount fluctuation and loss details consistency on R d The impact of
[0063] The proposed medical report or supporting evidence information layer risk score is R m ; Medical report or supporting evidence information layer risk score R m Based on medical report or supporting evidence information layer L m The file source verification parameter P in mr , file structure consistency parameter P mf to measure risk;
[0064] R m =g·P mr +h·P mf
[0065] Where R m It is used to reflect whether there are any anomalies in the source and structure of the certification documents; g and h are adjustment coefficients used to adjust the effect of document source and structure consistency on R m The impact of
[0066] Risk score R based on four information levels s , R e , R d , R m Perform weighted aggregation to obtain an overall risk score R to determine whether to trigger an early warning;
[0067] R=w1·R s +w2·R e +w3·R d +w4·R m
[0068] R is the overall risk score, which is used to determine the comprehensive risk of the application; w1, w2, w3, and w4 are the weights of each level, which are used to balance the contribution of different levels to the total score; if the risk exceeds the preset threshold R threshold When the system triggers an early warning, the application will be marked as high risk;
[0069] After calculating the overall risk score R, compare R with the system preset risk threshold R threshold Compare; if R>R threshold , an early warning is triggered and the application is marked as high risk.
[0070] Technical effects and advantages of the present invention:
[0071] 1. The present invention structures the insurance claims data through multi-level information division and labeling, analyzes the risk characteristics of each type of information in a targeted manner, realizes multi-dimensional accurate review, effectively improves the recognition accuracy of abnormal data, reduces the misjudgment rate, and thus improves the accuracy and efficiency of insurance data review;
[0072] 2. The system uses the setting of anomaly detection module to compare with historical data, identify anomalies in identity information and event descriptions, reduce the pressure of manual review, make the review process more convenient, and reduce the deviation caused by human intervention;
[0073] 3. This solution calculates the risk scores of each information level and generates an overall risk score through comprehensive analysis. Once the score exceeds the preset threshold, the system automatically triggers an early warning. This mechanism provides a real-time monitoring method for high-risk applications, ensuring that high-risk cases can enter the manual review stage in a timely manner, thereby improving the initiative and response speed of risk management.
[0074] 4. The system is designed with a feedback and model optimization module to feed back manual review results to the model for learning and optimization; by continuously adjusting the weights and thresholds of each parameter, the system can dynamically adapt to new data and risk characteristics, and gradually improve the accuracy of the review and the optimization capabilities of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0076] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0077] Refer to the instruction manual Figure 1 , an insurance data review system based on a large language model according to an embodiment of the present invention comprises a data preprocessing module, an information layer division and marking module, an anomaly detection module, a risk scoring and early warning module, and a feedback and model optimization module;
[0078] The data preprocessing module is used to obtain the text data of the claim application before the insurance claim is settled, and to clean and standardize the text data, delete redundant information, and obtain preliminary processed data;
[0079] The information layer division and labeling module divides and labels the preliminary processed data according to the information layer through the large language model. The information layer includes the identity feature information layer, the event description information layer, the loss description information layer, and the medical report or auxiliary certificate information layer;
[0080] The anomaly detection module compares the current application data with the historical application data through a model. If the identity feature information layer detection or the event description information layer detection is determined to be abnormal, the risk assessment is performed through the risk scoring and early warning module;
[0081] The risk scoring and early warning module generates risk scores based on analysis results, marks high-risk applications and triggers early warnings, and is used for manual review;
[0082] The feedback and model optimization module optimizes the model recognition rules by feeding back the results of manual review to the model to improve the audit accuracy. Optimizing the model recognition rules means adjusting and improving the model's judgment parameters, thresholds, and classification rules based on the feedback data from manual review to improve the model's performance in identifying anomalies, judging risks, and auditing accuracy. By optimizing the recognition rules, the model can more accurately adapt to new data features and risk patterns, thereby reducing misjudgments and missed judgments, and improving the intelligence and accuracy of the overall audit system.
[0083] The solution constructs an insurance data review process through five modules: data preprocessing, information layer division, anomaly detection, risk assessment and early warning, and feedback and optimization. The purpose is to analyze various types of information in claim applications in a hierarchical and detailed manner. First, the original data is cleaned and standardized to ensure data consistency and normalization. Then, the information is classified and marked through a large language model to provide a basis for separate analysis of the information content of each layer. Then, through the anomaly detection module, the current application is compared with historical data to find abnormal information in a targeted manner. Once an anomaly is found, a risk score is generated and an early warning is issued through the risk assessment module to facilitate manual review. Finally, through the feedback and optimization module, the model recognition rules are adjusted according to the results of manual review. This processing method helps to improve the accuracy and efficiency of recognition, reduce the misjudgment rate, and achieve efficient automated review and continuous improvement.
[0084] In addition, the large language model in the above solution is an artificial intelligence model based on natural language processing (NLP) technology, which can understand and analyze complex semantics in texts. The large language model can select pre-trained models such as BERT, ChatGLM, MOSS, or BELLE that are widely used in text analysis. The above models are also open source and suitable for local deployment and real-world processing. The above models have been trained on massive text data and have excellent context understanding, classification, and generation capabilities, so they are also suitable for insurance data review. The reason for using a large language model is that it can deeply analyze the natural language descriptions in the claim application (such as event descriptions, medical reports), identify potential anomalies or ambiguities in the text, especially some subtle contradictions or abnormal expressions. In addition, the use of a large language model can divide the claim text into layers, such as extracting identity information, accident details, loss conditions, etc., and converting complex text data into structured data, which is convenient for subsequent anomaly detection and risk assessment. In specific applications, the large language model receives standardized input text, first classifies and marks the information layer, then detects the rationality and consistency of each information layer by layer, and combines historical data to determine whether there are anomalies, providing refined automation support for the entire review process.
[0085] Before insurance claims are settled, the data preprocessing module obtains the text data of the claim application, cleans and standardizes the text data, deletes redundant information, and obtains preliminary processed data. The preliminary processed data set is D;
[0086] D={T i ∣T i =Clean(R i )andR i ∈Raw Data}
[0087] Where D represents the cleaned and standardized data set for subsequent model use; T i represents the i-th cleaned record; R i represents the i-th raw data; Raw Data represents raw data, i.e. text data without any cleaning, processing or standardization; the function Clean(·) is used to delete redundant information, standardize the text format, and standardize the data; the result of data preprocessing is a structured data set, which is used by the large language model of the information layer division and labeling module for classification and anomaly detection;
[0088] The information layer division and labeling module classifies the preliminary processed data set D through a large language model, and divides and labels the data into four information layers; the proposed identity feature information layer is L s , the event description information layer is L e , the loss description information layer is L d, medical report or supporting certificate information layer is L m ;
[0089] D={L s ,L e ,L d ,L m}
[0090] Where L s , L e , L d , L m They represent four information layers respectively; the large language model divides and marks each piece of data into the corresponding information layer, providing a basis for subsequent parameter calculation and risk assessment; in addition, the layered processing method is adopted to structure the complex claim application text according to the information type, so that the large language model can process different types of information (such as identity information, event description, loss description, and supporting certificates) respectively, thereby improving the accuracy and pertinence of data analysis. The advantage of layering is that it can refine the information processing process and analyze different information layers separately, so as to more accurately identify potential problems in subsequent anomaly detection and risk assessment. Through the large language model, its powerful natural language understanding ability can be used to perform semantic analysis on the text and automatically divide and mark the information layer, classify each piece of data into the corresponding category, and provide a structured data foundation for subsequent steps.
[0091] The judgment conditions of the identity feature information layer detection include: when the identity information consistency parameter is lower than the identity information consistency parameter threshold preset by the system and the identity authentication frequency parameter is higher than the identity authentication frequency parameter threshold preset by the system, it is judged as abnormal;
[0092] The judgment conditions of the event description information layer detection include: when the event description similarity parameter is higher than the event description similarity parameter threshold preset by the system and the event time consistency parameter is lower than the event time consistency parameter threshold preset by the system, it is judged as abnormal;
[0093] This anomaly judgment method detects anomalies in identity information and event descriptions based on specific parameters, and makes hierarchical judgments on the information by setting thresholds. If the identity information consistency parameter is low and the verification frequency is high, or the event description similarity parameter is high and the time consistency is low, potential fraud can be identified more accurately. The advantage of such judgment is that it can comprehensively consider the consistency and change frequency of information, prevent false applications or repeated descriptions, and improve the accuracy and efficiency of the review system.
[0094] Identity feature information layer L s Including evaluating the identity information consistency parameter P si , authentication frequency parameter P sf ;
[0095] Identity information consistency parameter P si Include name matching sub-parameter M n , ID number matching sub-parameter M c , Address matching sub-parameter M a ;
[0096] P si Used to measure the consistency of the identity information of the applicant's name, ID number and address submitted with the historical records, so as to identify whether there is a risk of inconsistency in the identity information;
[0097] P si =α·M n +β·M c +γ·M a
[0098] Where P si It is used to measure the overall matching degree of identity information; α, β, and γ are the consistency weights of name, ID number, and address, respectively, which are used to balance the impact of different information on the overall consistency;
[0099]
[0100]
[0101]
[0102] Where N is the total number of name matches in the history; Indicates the name weight of the i-th record, which is used to reflect the reliability of the record; Indicates whether the name of the i-th record matches, 1 if it matches, and 0 if it does not match; M n The closer it is to 1, the higher the name consistency; C is the total number of ID number matches in the historical records; Indicates the certificate number weight of the i-th record; Indicates whether the certificate number matches, if it matches, it is 1, if it does not match, it is 0; c The higher the value of, the better the consistency of the document number; A is the total number of address matches in the historical records; The weight of the address record; A flag indicating whether the address matches; high M a The value indicates a stronger consistency of the address.
[0103] Authentication frequency parameter P sf It is used to evaluate the frequency of identity information changes of applicants within a set time period and identify potential anomalies of frequent changes; the identity authentication frequency parameter P sf Including contact information change frequency F p , Address change frequency F a ;
[0104]
[0105] Where P sf Used to assess the frequency of changes in contact information and address; c and σ a are the weight coefficients of contact information and address, respectively, to adjust their impact on the change frequency. They can be set according to business needs. For example, if frequent changes in contact information are more risky than changes in address, then σ c Higher value; T is the length of the observation time window, which is used to calculate the standardized change frequency. It represents the total length of the evaluation period and can be a fixed time unit, such as one year, half a year, etc., so that the change frequency is associated with the time span; By p and F a Combine by weight and divide by the time window T to get the frequency of identity information change of the applicant in unit time; when P sf When the value is high, it indicates that many identity updates have occurred in a short period of time, which may indicate instability or abnormal behavior;
[0106]
[0107]
[0108] Where P represents the number of contact records; Indicates whether the i-th record has been changed, if changed it is 1, if not changed it is 0; T p represents the total number of contact method changes during the observation period; A represents the number of address records; Indicates the address change, if changed it is 1, if not changed it is 0; T a Indicates the total number of address changes during the observation period; where F p The higher the value, the greater the frequency of contact method changes; a The higher the value, the higher the frequency of address changes, which may be abnormal.
[0109] Event description information layer L e Including evaluating the event description similarity parameter P es , event time consistency parameter P et ;
[0110] Event description similarity parameter P es It is used to compare the similarity between the current event description and the historical record to help detect whether there is repeated description or fiction; the event description similarity parameter P es Including keyword similarity S k , sentence structure similarity S t ;
[0111] P es =λ·S k +μ·S t
[0112]
[0113]
[0114] Among them, λ and μ are adjustment coefficients, respectively controlling the keyword similarity S k Similarity with sentence structure S t P es If the keyword is more important, the λ value should be higher, otherwise if the sentence structure is more important, the μ value should be increased; K is the total number of keywords, that is, the number of keywords used to analyze similarity; w j is the weight of the jth keyword, which is used to control the influence of the keyword in the similarity calculation. If some keywords can better reflect the core content of the event, they can be given a higher weight; is the matching mark of the jth keyword, which is 1 if the keyword matches and 0 if it does not match. The similarity of the keywords is calculated by accumulating the matching status of each keyword. A high keyword match indicates a strong consistency in the event description. M is the total number of sentences, that is, the number of sentences used for similarity analysis. is the matching mark of the mth sentence. If the sentence structure matches, it is 1, otherwise it is 0. By accumulating the matching of the sentence structure, the similarity of the overall sentence structure is calculated. By calculating the similarity of keywords and sentence structure, the repeated or exaggerated components in the event description can be identified, providing a basis for judging abnormalities. The event description similarity P es The similarity of keywords and sentence structure is used to comprehensively measure the repetitiveness of descriptions. If the similarity of keywords and sentence structure is high, then P es The value will increase, indicating that the current event description may be too similar to the historical record and may be repeated or exaggerated;
[0115] Event time consistency parameter P et Used to measure the reasonableness between the time of event occurrence and the time of application submission, and help identify anomalies in time information. If the time consistency is insufficient, it may indicate false statements or abnormal circumstances;
[0116]
[0117] Where T s is the actual time of the event, T s Provided by the applicant, used to determine the starting time of the event; T e T is the application submission time, that is, the time when the applicant submits the claim application; maxThe maximum time deviation allowed defines a reasonable time range between the occurrence of an event and the submission of an application; the event time consistency parameter P et By calculating the difference between the event occurrence time and the application submission time, that is, |T s -T e |, and normalize the difference, that is, divide it by T max , thus obtaining the time consistency score; if P et A larger value indicates that there is a large difference between the event time and the application time, and there may be abnormal reporting.
[0118] Loss description information layer L d Including the floating parameter P of the estimated loss amount di , loss detail consistency parameter P da ;
[0119] Loss amount floating parameter P di Used to measure the difference between the loss amount of the current application and the historical record, and detect whether there is abnormal fluctuation in the amount; the loss detail consistency parameter P da Used to measure the consistency between current and historical loss details and detect whether there are differences in the detailed descriptions;
[0120]
[0121]
[0122] Among them, M current is the loss amount of the current claim; M average is the average loss amount of similar historical events; if P di A high value indicates that the loss amount fluctuates greatly, which may lead to abnormal reporting. i is the weight of the i-th loss detail; Indicates the matching of current and historical loss details; D is the total number of fields of loss details, that is, the number of loss details used for similarity comparison. For example, loss details can include "damaged part", "loss extent", "repair time", etc. Each loss detail is regarded as a field; P da It is used to evaluate the consistency between the loss details described in the current application and the descriptions of similar events in the historical records. The higher the value, the stronger the consistency.
[0123] Medical report or supporting evidence information layer L m Includes evaluation document source verification parameter P mr , file structure consistency parameter P mf ;
[0124] File origin verification parameter P mrUsed to detect the source consistency of the attached certification documents; file structure consistency parameter P mf Used to measure the similarity of file formats and structures;
[0125]
[0126]
[0127] Where R is the total number of document source fields, that is, the number of fields used to verify the document source. For example, the document source fields may include "issuing institution name", "doctor's signature", "date", etc. Each field is used for source consistency verification; w i is the weight of the i-th source field, which is used to adjust the impact of each source field on the overall source consistency. If a field, such as “issuing organization name”, is more critical than other fields, the weight of the field w i Set to a higher value; The matching status of the i-th source field is used to indicate the consistency between the current application document and the historical record on the i-th field. If there is no match, it is 0; file source verification parameter P mr The weighted consistency scores of each source field are accumulated and averaged to obtain the overall consistency score. Represents the weighted consistency of the i-th source field. If the field is consistent, the score is w i , if not, it is 0. Finally, the consistency scores of all fields are accumulated and divided by the total number of fields R to get the overall consistency score. The lower the P mr The value indicates that the document source is inconsistent and there is a risk of forgery;
[0128] Where F is the total number of file structure fields, that is, the number of fields used to analyze the consistency of the file structure. For example, the fields of the file structure may include "font type", "signature position", "paragraph format", etc. Each item is analyzed as a structure field; w j is the weight of the jth structure field, which is used to adjust the importance of each structure field in the overall structure consistency score. If a certain structure field, such as "signature position", is more important for file standardization, the weight of this field, w j Set to Higher; The jth structure field is used to indicate the consistency between the current file and the standard template in the jth structure field. If not consistent, it is 0; the file structure consistency parameter P mf The weighted consistency scores of each structural field are accumulated and averaged to obtain the overall structural consistency. The weighted consistency score of each field is If they are consistent, then the weight w is includedj , if not, it is 0. Finally, the consistency scores of all fields are added up and divided by the total number of fields F to get the overall file structure consistency score. mf The value may indicate that the file structure does not meet the standards and there is a risk of fraud.
[0129] The proposed identity feature information layer risk score is R s ; Identity feature information layer risk score R s Based on the identity feature information layer L s The identity information consistency parameter P in si and the authentication frequency parameter P sf To measure risk;
[0130] R s =a·P si +b·P sf
[0131] Where R s It is used to reflect whether there are abnormalities in the consistency and update frequency of the applicant's identity information; a and b are adjustment coefficients that determine P si and P sf R s The influence weight of . If identity consistency is more important, the value of a is set higher;
[0132] The risk score of the proposed event description information layer is R e ; Event description information layer risk score R e Based on the event description information layer L e The event description similarity parameter P in es and the event time consistency parameter P et To measure risk;
[0133] R e =c·P es +d·P et
[0134] Where R e It is used to reflect the repeatability and time consistency of event descriptions. c and d are adjustment coefficients, and their weights are set according to business needs. If the repeatability of event descriptions is more critical, c is set higher.
[0135] The proposed loss description information layer risk score is R d ; Risk score R of loss description information layer d Based on the loss description information layer L d The floating parameter P of the loss amount in di and loss detail consistency parameter P da To assess the consistency of loss amounts and detailed descriptions;
[0136] R d =e·P di +f·P da
[0137] Where R d It is used to reflect the fluctuation of loss amount and the consistency of loss details description; e and f are adjustment coefficients, which determine the impact of loss amount fluctuation and loss details consistency on R d The impact of
[0138] The proposed medical report or supporting evidence information layer risk score is R m ; Medical report or supporting evidence information layer risk score R m Based on medical report or supporting evidence information layer L m The file source verification parameter P in mr , file structure consistency parameter P mf to measure risk;
[0139] R m =g·P mr +h·P mf
[0140] Where R m It is used to reflect whether there are any anomalies in the source and structure of the certification documents; g and h are adjustment coefficients used to adjust the effect of document source and structure consistency on R m The impact of
[0141] Risk score R based on four information levels s , R e , R d , R m Perform weighted aggregation to obtain an overall risk score R to determine whether to trigger an early warning;
[0142] R=w1·R s +w2·R e +w3·R d +w4·R m
[0143] R is the overall risk score, which is used to determine the comprehensive risk of the application; w1, w2, w3, and w4 are the weights of each level, which are used to balance the contribution of different levels to the total score. The weights are set based on the relative importance of each information layer in fraud detection. For example, if the identity feature information layer is considered the most critical layer, w1 can be given a higher weight. If the risk exceeds the preset risk threshold R, the weight of the application will be increased. threshold When the system triggers an early warning, the application will be marked as high risk;
[0144] After calculating the overall risk score R, compare R with the system preset risk threshold R threshold Compare; if R>Rthreshold , an early warning is triggered and the application is marked as high risk; threshold It is the risk threshold preset by the system, which is used to determine whether further manual review of the application is required; if R exceeds the threshold, the system will automatically mark it as high risk and generate an early warning prompt to guide manual reviewers to focus on checking the application. Feedback and optimization Finally, the results of manual review are fed back to the model to optimize risk scoring and parameter settings. This feedback mechanism enables the system to dynamically adjust various weights and thresholds to continuously improve the accuracy and efficiency of review;
[0145] In addition, the scores of these four layers (identity feature information layer score, event description information layer score, loss description information layer score, and medical report or supporting certificate information layer score) can be used for risk assessment because they comprehensively cover the key information dimensions in the claim application, helping the system to evaluate the authenticity and rationality of the application from multiple angles. The identity feature information layer score evaluates the identity consistency and information update frequency of the applicant, and is mainly used to detect abnormal situations such as false identities or frequent identity changes; the event description information layer score helps to judge the authenticity of the application content by analyzing the similarity and time consistency of the event description, especially the risk of fictitious or exaggerated events; the loss description information layer score The separate rules assess the fluctuation of loss amount and the consistency of details to ensure that the claim amount and loss description are in line with the actual situation; the medical report or supporting certificate information layer score verifies the source of the document and the consistency of the structure to ensure the authenticity and completeness of the provided supporting documents. The selection of these four layers of scores can conduct a detailed review from the four core aspects of identity, incidents, losses and documents, effectively identify potential fraud, and improve the accuracy of risk judgment. The application of these scores in risk assessment generates an overall risk score by weighted aggregation of the scores at each layer. Whether the score exceeds the preset threshold is used to determine whether further review or warning triggering is required, thus realizing an intelligent risk management process.
[0146] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An insurance data review system based on a large language model, comprising a data preprocessing module, an information layer division and marking module, an anomaly detection module, a risk scoring and early warning module, and a feedback and model optimization module, characterized in that: The data preprocessing module is used to obtain the text data of the claim application before the insurance claim is settled, and to clean and standardize the text data, delete redundant information, and obtain preliminary processed data; The information layer division and labeling module divides and labels the preliminary processed data according to the information layer through the large language model. The information layer includes the identity feature information layer, the event description information layer, the loss description information layer, and the medical report or auxiliary certificate information layer; The anomaly detection module compares the current application data with the historical application data through a model. If the identity feature information layer detection or the event description information layer detection is determined to be abnormal, the risk assessment is performed through the risk scoring and early warning module; The risk scoring and early warning module generates risk scores based on analysis results, marks high-risk applications and triggers early warnings, and is used for manual review; The feedback and model optimization module feeds back the results of manual review to the model, optimizes the model recognition rules, and is used to improve audit accuracy.
2. The insurance data review system based on a large language model according to claim 1, characterized in that: Before insurance claims are settled, the data preprocessing module obtains the text data of the claim application, cleans and standardizes the text data, deletes redundant information, and obtains preliminary processed data. The preliminary processed data set is D; D={T i ∣T i =Clean(R i )andR i ∈Raw Data} Where D represents the cleaned and standardized dataset; T i represents the i-th cleaned record; R i represents the i-th raw data; Raw Data represents the raw data; the function Clean(·) is used to delete redundant information, standardize the text format, and standardize the data; the result of data preprocessing is a structured data set, which is used by the large language model of the information layer division and labeling module for classification and anomaly detection; The information layer division and labeling module classifies the preliminary processed data set D through a large language model, and divides and labels the data into four information layers; the proposed identity feature information layer is L s , the event description information layer is L e , the loss description information layer is L d , medical report or supporting certificate information layer is L m ; D={L s ,L e ,L d ,L m } Where L s , L e , L d , L m Represents four information layers respectively.
3. The insurance data review system based on a large language model according to claim 2, characterized in that: The judgment conditions of the identity feature information layer detection include: when the identity information consistency parameter is lower than the identity information consistency parameter threshold preset by the system and the identity authentication frequency parameter is higher than the identity authentication frequency parameter threshold preset by the system, it is judged as abnormal; The judgment conditions for event description information layer detection include: when the event description similarity parameter is higher than the event description similarity parameter threshold preset by the system and the event time consistency parameter is lower than the event time consistency parameter threshold preset by the system, it is judged as abnormal.
4. The insurance data review system based on a large language model according to claim 3, characterized in that: Identity feature information layer L s Including evaluating the identity information consistency parameter P si , authentication frequency parameter P sf ; Identity information consistency parameter P si Include name matching sub-parameter M n , ID number matching sub-parameter M c , Address matching sub-parameter M a ; P si Used to measure the consistency of the identity information of the applicant's name, ID number and address submitted with the historical records, so as to identify whether there is a risk of inconsistency in the identity information; P si =α·M n +β·M c +γ·M a Where P si It is used to measure the overall matching degree of identity information; α, β, and γ are the consistency weights of name, ID number, and address, respectively, which are used to balance the impact of different information on the overall consistency; Where N is the total number of name matches in the history; Indicates the name weight of the i-th record; Indicates whether the name of the i-th record matches; C is the total number of ID number matches in the historical records; Indicates the certificate number weight of the i-th record; Indicates whether the certificate number matches, 1 for a match and 0 for a mismatch; A is the total number of address matches in the history record; The weight of the address record; A flag indicating whether the address matches; Authentication frequency parameter P sf Including contact information change frequency F p , Address change frequency F a ; Where P sf Used to assess the frequency of changes in contact information and address; c and σ a are the weight coefficients of contact information and address, respectively, to adjust their influence on the change frequency; T is the length of the observation time window; Where P represents the number of contact records; Indicates whether the i-th record has changed; T p represents the total number of contact method changes during the observation period; A represents the number of address records; Indicates address change; T a Indicates the total number of address changes during the observation period.
5. The insurance data review system based on a large language model according to claim 4, characterized in that: Event description information layer L e Including evaluating the event description similarity parameter P es , event time consistency parameter P et ; Event description similarity parameter P es Used to compare the similarity between the current event description and the historical record; event description similarity parameter P es Including keyword similarity S k , sentence structure similarity S t ; P es =λ·S k +μ·S t Among them, λ and μ are adjustment coefficients, respectively controlling the keyword similarity S k Similarity with sentence structure S t P es The influence weight of K is the total number of keywords; w is j is the weight of the jth keyword; is the matching mark of the jth keyword; M is the total number of sentences; is the matching tag of the mth sentence; Event time consistency parameter P et Used to measure the reasonableness between the time when the event occurred and the time when the application was submitted; Where T s The actual time of the event; T e T is the application submission time, that is, the time when the applicant submits the claim application; max is the maximum allowed time deviation.
6. The insurance data review system based on a large language model according to claim 5, characterized in that: Loss description information layer L d Including the floating parameter P of the estimated loss amount di , loss detail consistency parameter P da ; Loss amount floating parameter P di Used to measure the difference between the loss amount of the current application and the historical record, and detect whether there is abnormal fluctuation in the amount; the loss detail consistency parameter P da Used to measure the consistency between current and historical loss details and detect whether there are differences in the detailed descriptions; Among them, M current is the loss amount of the current claim; M average is the average loss amount of similar historical events; w i is the weight of the i-th loss detail; Indicates the matching of current and historical loss details; D is the total number of fields in the loss details; Medical report or supporting evidence information layer L m Includes evaluation document source verification parameter P mr , file structure consistency parameter P mf ; File origin verification parameter P mr Used to detect the source consistency of the attached certification documents; file structure consistency parameter P mf Used to measure the similarity of file formats and structures; Where R is the total number of file source fields; w i is the weight of the i-th source field, which is used to adjust the impact of each source field on the overall source consistency; is the matching status of the i-th source field; Where F is the total number of file structure fields; w j is the weight of the jth structure field; It is the matching status of the jth structure field.
7. The insurance data review system based on a large language model according to claim 6, characterized in that: The proposed identity feature information layer risk score is R s ; Identity feature information layer risk score R s Based on the identity feature information layer L s The identity information consistency parameter P in si and the authentication frequency parameter P sf To measure risk; R s =a·P si +b·P sf Where R s It is used to reflect whether there are abnormalities in the consistency and update frequency of the applicant's identity information; a and b are adjustment coefficients that determine P si and P sf R s The weight of influence; The risk score of the proposed event description information layer is R e ; Event description information layer risk score R e Based on the event description information layer L e The event description similarity parameter P in es and the event time consistency parameter P et To measure risk; R e =c·P es +d·P et Where R e It is used to reflect the repeatability and time consistency of event description; c and d are adjustment coefficients; The proposed loss description information layer risk score is R d ; Risk score R of loss description information layer d Based on the loss description information layer L d The floating parameter P of the loss amount in di and loss detail consistency parameter P da To assess the consistency of loss amounts and detailed descriptions; R d =e·P di +f·P da Where R d It is used to reflect the fluctuation of loss amount and the consistency of loss details description; e and f are adjustment coefficients, which determine the impact of loss amount fluctuation and loss details consistency on R d The impact of The proposed medical report or supporting evidence information layer risk score is R m ; Medical report or supporting evidence information layer risk score R m Based on medical report or supporting evidence information layer L m The file source verification parameter P in mr , file structure consistency parameter P mf To measure risk; R m =g·P mr +h·P mf Where R m It is used to reflect whether there are any anomalies in the source and structure of the certification documents; g and h are adjustment coefficients used to adjust the effect of document source and structure consistency on R m The impact of Risk score R based on four information levels s , R e , R d , R m Perform weighted aggregation to obtain an overall risk score R to determine whether to trigger an early warning; R=w1·R s +w2·R e +w3·R d +w4·R m R is the overall risk score, which is used to determine the comprehensive risk of the application; w1, w2, w3, and w4 are the weights of each level, which are used to balance the contribution of different levels to the total score; if the risk exceeds the preset threshold R threshold When the system triggers an early warning, the application will be marked as high risk; After calculating the overall risk score R, compare R with the system preset risk threshold R threshold Compare; if R>R threshold , an early warning is triggered and the application is marked as high risk.
Citation Information
Cited By
Intelligent claim settlement auxiliary system based on artificial intelligence
CN121120266A