Insurance policy claim data screening method and device, electronic equipment and storage medium

CN122134472APending Publication Date: 2026-06-02PING AN TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-02-04
Publication Date
2026-06-02

Smart Images

  • Figure CN122134472A_ABST
    Figure CN122134472A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of policy claim data screening method and device, electronic equipment and storage medium, belong to artificial intelligence technical field, apply to financial scene.The method comprises: obtaining policy claim multi-modal data, wherein the policy claim multi-modal data includes policy claim event data and claim result data;Based on the pre-constructed claim decision model, the policy claim event data is predicted, and the model claim decision is obtained;Based on the claim result data and the model claim decision, causal effect evaluation is carried out, and the causal effect evaluation data is obtained;The positive and negative fact difference of the policy claim event data is calculated, and the positive and negative claim difference data is obtained;Based on the causal effect evaluation data and the positive and negative claim difference data, high-quality data screening is carried out on the policy claim multi-modal data.The quality of the screened policy claim data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applicable to the financial field, particularly to a method and apparatus for screening insurance policy claims data, an electronic device, and a storage medium. Background Technology

[0002] Policy claims data screening is used to select target data from a large amount of policy claims data that has a clear causal relationship with the business, has a significant impact on decision adjustments, and has high value for iterative optimization of the claims decision model. Taking the auto insurance claims scenario as an example, by screening a large amount of auto insurance claims data, it is possible to select case data that has a high degree of fit between the model decision and the anti-fraud objective, and whose claims costs or customer complaint rates change significantly after adjusting the claims decision.

[0003] Currently, traditional policy claim data screening methods randomly select from numerous policy claim data points and then apply the selected policy claim data to the model optimization process. However, the policy claim data selected by this screening method is often disconnected from business objectives. Therefore, how to improve the quality of the screened policy claim data has become an urgent problem to be solved. Summary of the Invention

[0004] To address the technical problem of improving the quality of filtered policy claim data, the main objective of this application is to propose a policy claim data filtering method, apparatus, electronic device, and storage medium, which aims to improve the quality of the filtered policy claim data.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for filtering policy claims data, the method comprising:

[0006] Acquire multimodal data on policy claims, wherein the multimodal data on policy claims includes policy claim event data and claim result data; Based on a pre-built claims decision model, claims prediction is performed on the policy claims event data to obtain model claims decisions; Based on the claims result data and the claims decision of the model, a causal effect assessment is performed to obtain causal effect assessment data. The positive and negative factual differences are calculated on the policy claim event data to obtain positive and negative claim difference data; Based on the causal effect assessment data and the positive and negative claims difference data, high-quality data screening is performed on the multimodal data of the policy claims.

[0007] In some embodiments, the step of conducting a causal effect assessment based on the claims outcome data and the model claims decision to obtain causal effect assessment data includes: The model's claims decision is then subjected to parameter intervention to obtain positive parameter intervention schemes and negative parameter intervention schemes; Based on the aforementioned positive intervention scheme, the expected value of the claims result data is calculated to obtain the expected value of the positive intervention scheme. Based on the aforementioned parameter reverse intervention scheme, the mathematical expectation of the claims result data is calculated to obtain the mathematical expectation data of parameter reverse intervention; Based on the expected mathematical data of positive intervention and the expected mathematical data of negative intervention of the parameters, the causal effect is calculated to obtain the causal effect assessment data.

[0008] In some embodiments, the step of calculating the difference between positive and negative facts in the policy claim event data to obtain the difference data between positive and negative claims includes: Counterfactual simulations were performed on the policy claim event data to obtain counterfactual claim results; Based on the counterfactual claim results and the claim result data, the difference in claim results is calculated to obtain the difference data between the positive and negative claims.

[0009] In some embodiments, performing counterfactual simulation on the policy claim event data to obtain counterfactual claim results includes: Decision variables are extracted from the multimodal data of the policy claims to obtain the claims decision; Based on the aforementioned claims decision, construct counterfactual input samples; Based on the counterfactual input sample, a claims simulation is performed to obtain the counterfactual claims result.

[0010] In some embodiments, the step of calculating the difference in claims results based on the counterfactual claims result and the claims result data to obtain the difference data between the positive and negative claims includes: Feature extraction is performed on the counterfactual claim results to obtain counterfactual claim features; Feature extraction is performed on the claims result data to obtain the actual claims features; The difference between the counterfactual claim features and the factual claim features is calculated using the same dimension feature difference to obtain the positive and negative claim difference data.

[0011] In some embodiments, the high-quality data screening of the multimodal policy claims data based on the causal effect assessment data and the positive and negative claims difference data includes: Based on the causal effect assessment data, a consistency assessment is performed on the model's claims decision and the claims result data to obtain a data consistency score. Based on the data consistency score, the multimodal data of the policy claims is initially screened to obtain candidate policy claims data; Based on the positive and negative claims data, the candidate policy claims data are screened twice to obtain high-quality policy claims data.

[0012] In some embodiments, the step of predicting claims based on the policy claim event data using a pre-built claims decision model to obtain a model-based claims decision includes: Multimodal feature extraction is performed on the policy claim event data to obtain the multimodal features of the policy claim events; The multimodal features of the policy claim event are fused to obtain the standard features of the policy claim event; Based on the claims decision model, feature mapping is performed on the standard features of the policy claims event to obtain the model claims decision.

[0013] To achieve the above objectives, a second aspect of this application provides a policy claims data filtering device, the device comprising: The insured object data acquisition module is used to acquire the insured object of the target policy and acquire the target object data of the insured object; The multidimensional assessment data acquisition module is used to filter the target object data based on preset target risk assessment dimensions to obtain target object assessment data for each target risk assessment dimension; wherein, the target risk assessment dimension includes at least one of the following: regional risk dimension, object ontology risk dimension, Internet of Things device risk dimension, and user behavior risk dimension. The feature extraction module is used to extract features from the assessment data of each target object to obtain the target object assessment features of each target risk assessment dimension. The policy risk assessment module is used to assess the policy risk of the target policy based on the assessment characteristics of each target object, and obtain the policy risk level. The risk warning module is used to perform policy risk warning operations based on the policy risk level.

[0014] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0015] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of the first aspect described above.

[0016] The policy claims data screening method, apparatus, electronic device, and storage medium proposed in this application first systematically integrates multimodal policy claims data. Based on a pre-built claims decision model, it predicts claims for policy claims events within the multimodal data, generating accurate model-based claims decisions. Then, it combines these model-based decisions with claims outcome data from the multimodal data to assess causal effects, obtaining causal effect assessment data. Simultaneously, it calculates the difference between positive and negative facts in the policy claims event data, obtaining positive and negative claims difference data. Finally, it integrates the causal effect assessment data and the positive and negative facts... Anti-claim discrepancy data enables high-quality data filtering of multimodal policy claims data. It effectively removes redundant information and noise interference from multimodal policy claims data, ensuring that the filtered data has real business causal relationships and significant optimization value. Through dual verification of causal effect assessment and positive and negative fact difference calculation, it improves the scientificity and accuracy of policy claims data filtering, and provides high-quality data support for the iterative optimization of claims decision models. In turn, it can enhance the claims decision models' ability to understand claims business logic and improve decision accuracy, thereby helping to improve the processing efficiency and rationality of policy claims business. Attached Figure Description

[0017] Figure 1 This is a flowchart of the policy claims data filtering method provided in the embodiments of this application; Figure 2 yes Figure 1 The flowchart of step S102 in the document; Figure 3 yes Figure 1 The flowchart of step S103 in the process; Figure 4 yes Figure 3 The flowchart of step S301 in the process; Figure 5 yes Figure 3 The flowchart of step S302 in the document; Figure 6 yes Figure 1 The flowchart of step S104 in the process; Figure 7 yes Figure 1 The flowchart of step S105 in the process; Figure 8 This is a schematic diagram of the structure of the policy claims data filtering device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0022] This application provides a method and apparatus for filtering policy claim data, an electronic device, and a storage medium, which aim to improve the quality of the filtered policy claim data.

[0023] The policy claim data filtering method, device, electronic equipment, and storage medium provided in this application are specifically described through the following embodiments. First, the policy claim data filtering method in this application embodiment is described.

[0024] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0025] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0026] The policy claim data filtering method provided in this application relates to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the policy claim data filtering method, but is not limited to the above forms.

[0027] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions performed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can be located on local and remote computer storage media, including storage devices.

[0028] Figure 1This is an optional flowchart of the policy claims data filtering method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.

[0029] Step S101: Obtain multimodal data of policy claims, which includes policy claim event data and claim result data; Step S102: Based on the pre-built claims decision model, perform claims prediction on the policy claims event data to obtain the model claims decision; Step S103: Based on the claims result data and model claims decision, conduct a causal effect assessment to obtain causal effect assessment data; Step S104: Calculate the difference between positive and negative facts in the policy claim event data to obtain the difference data between positive and negative claims. Step S105: Based on the causal effect assessment data and the difference data between positive and negative claims, high-quality data screening is performed on the multimodal data of policy claims.

[0030] Steps S101 to S105 as illustrated in this embodiment first involve systematically integrating multimodal policy claims data. Based on a pre-built claims decision model, claims prediction is performed on policy claims event data within the multimodal data, generating accurate model-based claims decisions. Then, the model-based claims decisions are combined with claims outcome data from the multimodal data to conduct a causal effect assessment, yielding causal effect assessment data. Simultaneously, the positive and negative fact differences in the policy claims event data are calculated to obtain positive and negative claims difference data. Finally, the causal effect assessment data and the positive and negative claims difference data are integrated. Heterogeneous data filtering enables high-quality screening of multimodal policy claims data. It effectively removes redundant information and noise interference from multimodal policy claims data, ensuring that the filtered data has real business causal relationships and significant optimization value. Through dual verification of causal effect assessment and positive and negative fact difference calculation, it improves the scientificity and accuracy of policy claims data screening. It also provides high-quality data support for the iterative optimization of claims decision models, thereby enhancing the claims decision models' ability to understand claims business logic and improve decision accuracy, thus helping to improve the processing efficiency and rationality of policy claims business.

[0031] In step S101 of some embodiments, multimodal data of policy claims refers to a collection of various types of data collected in the policy claims business. In the financial insurance scenario, multimodal data of policy claims may include customer identity information, policy terms text, medical diagnosis report, medical expense invoice images, claims case processing records, etc.

[0032] Policy claim event data refers to information related to the basic events that trigger policy claims. In the context of financial insurance, policy claim event data can include claim report details, time and location of the incident, explanation of the cause of the accident, description of the damage to the insured property, and relevant liability determination materials.

[0033] Claims outcome data refers to the result information generated after the policy claims case has been processed. In the financial insurance scenario, claims outcome data can include the final compensation amount, the determination of whether or not compensation is paid, the claims processing time, customer satisfaction feedback, and the results of claims dispute resolution, etc.

[0034] This application embodiment can integrate policy claim event data and claim result data through a multi-channel data collection mechanism within the enterprise's policy claim business system, thereby merging the policy claim event data and claim result data into a complete multimodal data set. It is important to understand that the sources of policy claim event data and claim result data can include materials submitted by customers through online channels, information entered during offline business processing, and process data automatically recorded by the system. In the auto insurance claim scenario, data such as customer identity information, policy terms details, accident scene images, scanned copies of repair expense invoices, and voice-to-text records of the report can be collected to form policy claim event data, as well as data such as the final compensation amount, claim processing time, and customer feedback results to form claim result data. In the health insurance claim scenario, data such as inpatient medical records, examination imaging data, and medication list forms can be collected to form policy claim event data, as well as data such as reimbursement amount and approval status to form policy claim event data.

[0035] In step S102 of some embodiments, the claims decision model refers to a model used to analyze policy claims and output decision recommendations. In the financial insurance scenario, the claims decision model can be an intelligent algorithm model used to determine the authenticity of claims, assess the scope of compensation liability, and estimate the reasonable compensation amount.

[0036] Model-based claims decision refers to the decision result obtained by the claims decision model after analyzing policy claims cases. In the financial insurance scenario, model-based claims decision can be an opinion on whether to agree to pay compensation, a specific compensation amount suggestion, a prompt on whether manual review is required, and a result of the claims risk level assessment.

[0037] This application embodiment extracts and fuses multimodal features from policy claim event data to obtain unified standard features. Then, a pre-trained claim decision model is used to map these standard features to obtain model claim decisions for claim cases in the policy claim event data.

[0038] For details, please refer to Figure 2In some embodiments, step S102 includes, but is not limited to, steps S201 to S203: Step S201: Extract multimodal features from the policy claim event data to obtain the multimodal features of the policy claim event; Step S202: Perform feature fusion on the multimodal features of the policy claim event to obtain the standard features of the policy claim event; Step S203: Based on the claims decision model, perform feature mapping on the standard features of policy claims events to obtain the model claims decision.

[0039] In step S201 of some embodiments, the multimodal features of the policy claim event refer to a set of various types of features extracted through multimodal features. In the car insurance claim scenario, the multimodal features of the policy claim event can be the semantic features of the accident description in the car insurance claim, the visual features of the scene photos, the structured features of the policy terms, etc.

[0040] This application first categorizes policy claim event data into text, image, and form types based on data type. Then, different feature extraction techniques are employed for different data types to obtain multimodal features of policy claim events. For example, text-based policy claim event data, such as claim report descriptions and policy terms, can be parsed using natural language processing to extract key features and obtain text features. Image-based policy claim event data, such as accident scene photos and scanned expense receipts, can be identified using computer vision technology to extract visual information and features and obtain image features. Form-based policy claim event data, such as insurance information and damage details, can have its structured fields directly extracted as form features. Finally, by integrating the features of each modality, multimodal features of policy claim events can be obtained. For example, in the car insurance claims scenario, semantic features such as "rear-end collision accident, front bumper damage" can be extracted from the customer's report text, visual features such as "damage level, vehicle model" can be extracted from the accident photos, and structured features such as "100,000 yuan coverage amount, 2-year insurance period" can be extracted from the insurance application form. Finally, by integrating the features of the aforementioned modalities, a multimodal feature of the car insurance policy claims event can be formed.

[0041] In step S202 of some embodiments, the standard features of the policy claim event can be a fixed-dimensional real number vector that can comprehensively and systematically represent the core information of the policy claim event.

[0042] The embodiments of this application can first replace various features in the multimodal features of the insurance claim event with a compatible format, and then integrate various features such as text, images, and forms into a standardized feature vector of a unified dimension through methods such as splicing, weighted fusion or deep learning fusion network, thereby achieving the elimination of heterogeneity of different modal features.

[0043] In step S203 of some embodiments, the pre-built claims decision model has been trained with massive amounts of historical financial claims data and possesses the ability to learn the correlation between features and decision results. Therefore, after inputting the standard features of the policy claims event into the claims decision model in vector form, the claims decision model can analyze the matching relationship between the standard features of the policy claims event and the compensation rules and risk levels through the internal algorithm logic learned during model training, and complete the mapping from features to decision results to obtain the model claims decision. For example, in the corporate property insurance claims scenario, after analyzing the standard features of the corporate property insurance claims event, the claims decision model can output specific model claims decisions such as "agree to pay, suggest a compensation amount of 50,000 yuan, no manual review required".

[0044] Steps S201 to S203 as illustrated in this embodiment first extract multimodal features from the policy claim event data to obtain multimodal features of the policy claim event. This can comprehensively capture the core value information of policy claim event data of different modalities and avoid the limitations of single modal features. Secondly, feature fusion is performed on the multimodal features of the policy claim event to obtain standard features of the policy claim event. This can eliminate the processing obstacles caused by heterogeneous data and improve the standardization and completeness of model input. Finally, feature mapping is performed on the standard features of the policy claim event through the claim decision model, which can improve the accuracy and reliability of the obtained model claim decision and provide solid support for causal effect assessment, high-quality data screening and model iterative optimization.

[0045] In step S103 of some embodiments, the causal effect assessment data may be specific numerical values ​​or indicators of the impact of model claims decisions on reducing claims disputes, increasing business revenue, and improving risk control accuracy.

[0046] This application embodiment, by performing positive and negative bidirectional intervention on the parameters in the model's claims decision-making, can obtain two schemes after parameter intervention. Furthermore, the mathematical expectation of the claims result data under the two schemes is calculated respectively. Then, by calculating the difference between the two types of mathematical expectations, causal effect evaluation data that quantifies the true correlation between the model's claims decision and the claims result can be obtained.

[0047] For details, please refer to Figure 3 In some embodiments, step S103 includes, but is not limited to, steps S301 to S304: Step S301: Parameter intervention is applied to the model's claims decision to obtain a positive parameter intervention plan and a negative parameter intervention plan; Step S302: Based on the parameter positive intervention plan, perform mathematical expectation calculation on the claims result data to obtain the mathematical expectation data of parameter positive intervention; Step S303: Based on the parameter reverse intervention scheme, perform mathematical expectation calculation on the claims result data to obtain the parameter reverse intervention mathematical expectation data; Step S304: Based on the expected mathematical data of positive intervention and the expected mathematical data of negative intervention, calculate the causal effect to obtain the causal effect assessment data.

[0048] In step S301 of some embodiments, the parameter positive intervention scheme refers to the intervention setting that forces the adoption of the model's claims decision, such as a scheme that executes the compensation recommendation and liability determination results output by the model for all similar claims.

[0049] A parameter-reverse intervention scheme refers to an intervention setting that forces the rejection of the model's claims decision and executes the opposite logic. For example, a scheme that rejects the model's compensation recommendations for all similar claims and overturns the model's liability determination results.

[0050] This application embodiment can first clarify the core decision dimensions of the model's claims decision, such as whether to pay, the range of payout amount, and the result of liability determination. Then, based on these core decision dimensions, two types of intervention schemes can be constructed. Specifically, the parameter positive intervention scheme can be set to enforce the model's claims decision for all similar claims without adding any manual adjustments. The parameter negative intervention scheme can be set to enforce logic that is completely opposite to the model's claims decision for all similar claims. For example, if the model's claims decision recommends payment, the parameter negative intervention scheme can be to refuse payment. If the model's claims decision recommends payment of 50,000 yuan, the parameter negative intervention scheme can be to process it as 0 yuan or other opposite rules.

[0051] In step S302 of some embodiments, the expected mathematical data of positive parameter intervention refers to the average level of claims results data of similar claims cases when the positive parameter intervention plan is implemented. In the financial insurance scenario, the expected mathematical data of positive parameter intervention can be average claims efficiency, average claims cost, etc.

[0052] This application embodiment can collect claims outcome data such as payout amount, processing time, and customer complaint rate after implementing a positive intervention plan for similar claims cases. Furthermore, by calculating the average value of these claims outcome data—that is, the mathematical expectation—the expected value of the positive intervention can be obtained. For example, in a health insurance claims scenario, if 1000 inpatient claims cases with positive intervention are collected, and the average payout amount for these inpatient claims is 8000 yuan, the average processing time is 3 working days, and the average complaint rate is 2%, then these average values ​​are the expected value of the positive intervention for the corresponding indicators.

[0053] In step S303 of some embodiments, the mathematical expectation data of parameter reverse intervention refers to the average level of claim results data of similar claims cases when the parameter reverse intervention scheme is executed. In the financial insurance scenario, the mathematical expectation data of parameter reverse intervention may be the average dispute occurrence rate, average customer churn rate, etc. of similar cases after the rejection model claims decision.

[0054] This application embodiment can use statistical standards consistent with those used in calculating the expected mathematical data for positive parameter intervention to collect claims outcome data from similar claims cases implementing a negative parameter intervention plan. Then, by calculating the average value for the same core business indicators, the expected mathematical data for negative parameter intervention can be obtained. It is important to note that during the calculation of the expected mathematical data for negative parameter intervention, it is crucial to ensure that the number of data samples and the case types are consistent with the calculation scenario for the expected mathematical data for positive parameter intervention to guarantee the comparability of the two types of expected mathematical data.

[0055] In step S304 of some embodiments, the difference between the expected mathematical data of positive intervention and the expected mathematical data of negative intervention is calculated for the same business indicator. The difference result is the causal effect assessment data for that business indicator. It should be noted that if the expected mathematical data of positive intervention is better than the expected mathematical data of negative intervention, the difference between the two is usually positive. In this case, the causal effect assessment data indicates that the model's claims decision has a positive causal impact on the business indicator. Similarly, if the difference between the expected mathematical data of positive intervention and the expected mathematical data of negative intervention is negative, the causal effect assessment data indicates that the model's claims decision has a negative causal impact on the business indicator. For example, in a health insurance claims scenario, the causal effect assessment data for the payout amount is 5000 yuan, the causal effect assessment data for processing time is -4 working days, and the causal effect assessment data for the complaint rate is -13%. It can be clearly seen that the model's claims decision has a positive causal impact on the payout amount and a negative causal impact on the processing time and complaint rate.

[0056] Steps S301 to S304 as illustrated in this embodiment of the application, by constructing two types of intervention schemes—positive and negative—can ensure a rigorous comparative basis for evaluating the causal effect of the model's claims decision. Then, by calculating the mathematical expectation of the claims result data under the two types of intervention schemes, the interference of accidental factors in individual cases can be effectively eliminated, accurately reflecting the average business performance of similar cases. Finally, the causal effect evaluation data is obtained through difference calculation, realizing a quantitative representation of the true correlation between the model's decision and the claims result.

[0057] In step S104 of some embodiments, the difference data between positive and negative claims can be the difference in payout amount, claims processing cycle, customer retention rate, and risk loss amount under different decision assumptions.

[0058] This application embodiment can generate virtual counterfactual claim results by performing counterfactual simulation on policy claim event data, and then calculate the difference in claim results by comparing the counterfactual claim results with the actual claim results data, and finally obtain positive and negative claim difference data that quantifies the difference between the two types of results.

[0059] For details, please refer to Figure 4 In some embodiments, step S104 includes, but is not limited to, steps S201 to S402: Step S401: Perform counterfactual simulation on the policy claim event data to obtain counterfactual claim results; Step S402: Based on the counterfactual claim results and claim result data, calculate the difference in claim results to obtain the difference data between the positive and negative claims.

[0060] In step S401 of some embodiments, the counterfactual claim result refers to the virtual claim result generated through counterfactual simulation. For example, in the financial insurance scenario, the counterfactual claim result can be a virtual result such as the customer feedback status after simulating "no compensation" or the dispute occurrence rate after simulating "adjustment of compensation ratio".

[0061] This application embodiment extracts decision variables related to claims decision from multimodal data of insurance policy claims, then constructs counterfactual input samples that conform to the opposite decision logic based on these decision variables, and finally outputs virtual counterfactual claims results by performing claims simulation on the counterfactual input samples.

[0062] For details, please refer to Figure 5 In some embodiments, step S401 includes, but is not limited to, steps S501 to S503: Step S501: Extract decision variables from the multimodal data of policy claims to obtain the claims decision; Step S502: Based on the claims decision, construct counterfactual input samples; Step S503: Based on the counterfactual input sample, perform a claims simulation to obtain the counterfactual claims result.

[0063] In step S501 of some embodiments, the claims decision refers to the core parameters extracted from the multimodal data of the policy claims that directly determine the direction of the claims decision. The core parameters may be specific values ​​of key decision dimensions such as whether to pay, the payout ratio, and the liability determination result.

[0064] This application embodiment can, based on the core decision-making logic of claims business, filter out key parameters that directly affect the claims outcome from multimodal data of policy claims as claims decisions. It should be noted that the claims decision must have a clear decision orientation, such as whether to pay, the range of payout amount, whether manual review is required, etc., to ensure that it can accurately reflect the claims direction of multimodal data of policy claims.

[0065] In step S502 of some embodiments, the counterfactual input sample refers to virtual input data that retains the core basic features of the multimodal data of the policy claims and only replaces the relevant variables of the claims decision with the opposite logic. In the financial insurance scenario, the counterfactual input sample can be a sample that includes all other information of the original case after the "agree to pay" decision is adjusted to "reject pay".

[0066] This application embodiment can replace the extracted claims decision with a completely opposite decision logic while keeping all other basic features in the multimodal claims data of the policy unchanged, except for the decision variables, to form a counterfactual input sample. It is important to note that during the construction of the counterfactual input sample, it is crucial to ensure that the customer information, policy details, accident details, and other information in the counterfactual input sample are consistent with the real cases in the multimodal claims data of the policy, differing only in the decision direction. This ensures the validity of the subsequent simulation results comparison. For example, in a car insurance claims scenario, when the claims decision is "pay 4000 yuan," the parameters related to the claims decision in the counterfactual input sample could be "no payment," while all other features remain unchanged.

[0067] In step S503 of some embodiments, the constructed counterfactual input sample is input into a counterfactual simulation model that has been trained with massive amounts of historical financial claims data and has learned the correlation between decision logic and case outcome. The counterfactual simulation model can comprehensively analyze the basic features and opposite decision variables in the counterfactual input sample, deduce the possible trends of the case corresponding to the multimodal data of the policy claim under the opposite decision variables, and finally obtain the counterfactual claim result.

[0068] Steps S501 to S503 as illustrated in this embodiment of the application accurately extract core decision variables related to claims decisions from multimodal data of policy claims, ensuring that the core direction of counterfactual simulation is highly consistent with the business decision-making logic. Then, by retaining the basic characteristics of real cases and only adjusting decision variables, counterfactual input samples are constructed, ensuring the comparability of simulation results with real results. Finally, professional claims simulation is performed based on counterfactual input samples, which can generate reliable counterfactual claims results. This not only improves the accuracy and rationality of counterfactual simulation, but also provides high-quality virtual result basis for calculating the difference between positive and negative facts, helping to accurately identify data samples with high value for model optimization, thereby supporting the iterative upgrade of the claims decision model.

[0069] In step S402 of some embodiments, features are first extracted from the counterfactual claim results and claim results data to obtain the corresponding counterfactual claim features and factual claim features. Then, the difference is calculated for the same dimension indicators of the two types of features to finally obtain positive and negative claim difference data that quantifies the difference between the two types of results.

[0070] For details, please refer to Figure 6 In some embodiments, step S402 includes, but is not limited to, steps S601 to S603: Step S601: Extract features from the counterfactual claim results to obtain counterfactual claim features; Step S602: Extract features from the claims result data to obtain the actual claims features; Step S603: Calculate the difference in features of the same dimension for the counterfactual claim features and the factual claim features to obtain the difference data between the positive and negative claims.

[0071] In step S601 of some embodiments, counterfactual claims features refer to core business indicators extracted from counterfactual claims results. Counterfactual claims features may include virtual payout amount, virtual complaint rate, virtual risk control effect, etc.

[0072] This application embodiment can select key indicators with business significance from counterfactual claim results as counterfactual claim features based on the key focus dimensions of the claims business. These key indicators include, but are not limited to, virtual compensation amount, virtual approval rate, virtual customer complaint rate, and virtual processing time. For example, in a car insurance claim scenario, if the counterfactual claim result is "no compensation for this case, customer complaint rate 10%, dispute processing time 5 working days," then "virtual compensation amount 0 yuan, virtual complaint rate 10%, virtual processing time 5 working days" can be extracted as counterfactual claim features.

[0073] In step S602 of some embodiments, the factual claims feature refers to the core business indicators extracted from the claims result data. The factual claims feature may be the actual amount of compensation paid, the actual customer satisfaction, the actual dispute occurrence rate, etc.

[0074] This application embodiment can employ an indicator system and screening criteria consistent with the counterfactual claims feature extraction method to extract corresponding core indicators from claims result data. For example, for the aforementioned auto insurance case, if the actual claims result data is "compensation of 2,500 yuan, no customer complaints, and processing time of 2 working days," then "actual compensation amount of 2,500 yuan, actual complaint rate of 0%, and actual processing time of 2 working days" can be extracted as factual claims features.

[0075] In step S603 of some embodiments, for counterfactual claim features and factual claim features of the same dimension, the difference can be calculated according to a unified calculation rule. Usually, the value of the factual claim feature is subtracted from the value of the counterfactual claim feature, or vice versa, so as to ensure that the direction of the difference can clearly reflect the impact of the decision.

[0076] Steps S601 to S603 as illustrated in this application embodiment first extract the core features of the counterfactual claims results and claims results data respectively, ensuring that the comparison of the two types of results has a clear business orientation and a unified statistical standard. Then, by calculating the difference of features in the same dimension, the degree of difference of claims results under different decision-making logics is accurately quantified, avoiding the distortion of difference assessment caused by dimension confusion. This not only ensures the accuracy and comparability of the difference data between positive and negative claims, but also provides an accurate quantitative basis for high-quality data screening.

[0077] Steps S401 to S402 shown in the embodiments of this application overcome the limitations of relying solely on real observation data through counterfactual simulation, supplementing the virtual claims results under different decision-making logics, providing a more comprehensive basis for decision impact assessment, and further quantifying the difference between the results under real and counterfactual decisions through precise calculation of claims result differences, clearly presenting the degree of impact of decision adjustments on business indicators, and ensuring the authenticity and comprehensiveness of the difference data.

[0078] In step S105 of some embodiments, the consistency assessment of the model claims decision and claims result data is first performed based on the above-mentioned causal effect assessment data to obtain a data consistency score. Then, the initial screening of the multimodal data of policy claims is completed based on the data consistency score to obtain candidate policy claims data. Then, the candidate data is screened a second time based on the positive and negative claims difference data to finally obtain high-quality policy claims data that meet the needs of business optimization.

[0079] For details, please refer to Figure 7In some embodiments, step S105 includes, but is not limited to, steps S701 to S703: Step S701: Based on the causal effect assessment data, conduct a consistency assessment on the model claims decision and claims result data to obtain a data consistency score; Step S702: Based on the data consistency score, perform preliminary screening of the multimodal data of policy claims to obtain candidate policy claims data; Step S703: Based on the difference between positive and negative claims data, perform a second screening of candidate policy claims data to obtain high-quality policy claims data.

[0080] In steps S701 and S702 of some embodiments, the data consistency score refers to a quantitative indicator obtained through data consistency assessment. The data consistency score can be a score reflecting the degree of fit between the claims result data and business objectives. If the causal effect assessment data shows that the model claims decision has a significant positive impact on the business indicators of the claims business, then the higher the data consistency score, the more the claims result data of the multimodal claims data of the policy meets business needs. If the causal effect assessment data shows that the model claims decision has a negative impact on the business indicators of the claims business, then the lower the data consistency score, the more the claims result data of the multimodal claims data of the policy meets business needs.

[0081] Candidate policy claims data refers to multimodal policy claims data that are retained after initial data screening and whose claims settlement plans have a positive impact on the policy claims event.

[0082] This application embodiment can determine the direction of influence of model claims decisions on business indicators of claims business based on the positive or negative value of causal effect assessment data. Furthermore, if the causal effect assessment data shows that model claims decisions have a significant positive impact on business indicators of claims business, and the parameters of multiple indicators in the claims result data are consistent with those in the model claims decision, it indicates that the claims decision in the claims result data will also have a positive impact on business indicators of claims business. In this case, the higher the data consistency score, the more likely the multimodal data of policy claims corresponding to the claims result data should be retained, thereby forming candidate policy claims data. If the causal effect assessment data shows that model claims decisions have a negative impact on business indicators of claims business, and the parameters of multiple indicators in the claims result data are consistent with those in the model claims decision, it indicates that the claims decision in the claims result data will also have a negative impact on business indicators of claims business. In this case, the lower the data consistency score, the more likely the multimodal data of policy claims corresponding to the claims result data should be retained, thereby forming candidate policy claims data.

[0083] In step S703 of some embodiments, high-quality policy claims data refers to high-value data obtained after initial and secondary screening. In the financial insurance scenario, high-quality policy claims data can be data with clear business causal relationships, significant impact on decision adjustments, and the ability to assist in model iteration and optimization.

[0084] This application embodiment sets a screening threshold for positive and negative claims data. For core difference indicators such as payout amount, customer complaint rate, and processing time, candidate data with difference values ​​exceeding the threshold are retained, thus obtaining high-quality policy claims data. High-quality policy claims data obtained through this method indicates a significant change in claims outcomes after decision adjustments, making it more valuable for claims decision-making models to learn the impact logic of different decisions and improve decision accuracy.

[0085] Steps S701 to S703 as illustrated in this application embodiment first perform consistency assessment and initial data screening on the model's claims decision and claims result data based on causal effect evaluation data. This effectively eliminates noisy data that has a negative impact on business, ensuring that the candidate policy claims data is highly consistent with business objectives. Then, based on the difference between positive and negative claims data, the candidate policy claims data is screened a second time, prioritizing the retention of data that has a significant impact on decision adjustments and high value for model optimization. This not only ensures the business adaptability of the screened data but also highlights the optimization value of the data, significantly improving the accuracy and efficiency of high-quality data screening. It provides accurate and high-value data support for the iterative upgrade of the claims decision model, helping the model to continuously optimize its decision logic.

[0086] This application first systematically integrates multimodal data on policy claims, and relies on a pre-built claims decision model to predict claims events within the multimodal data. This generates accurate model-based claims decisions. Next, the model-based claims decisions are combined with claims outcome data from the multimodal data to conduct a causal effect assessment, yielding causal effect assessment data. Simultaneously, the differences between positive and negative facts in the policy claims event data are calculated, resulting in positive and negative claims difference data. Finally, by merging the causal effect assessment data and the positive and negative claims difference data, a comprehensive assessment of policy claims can be achieved. High-quality data screening of multimodal claims data effectively removes redundant information and noise interference from policy claims data, ensuring that the screened data has real business causal relationships and significant optimization value. Through dual verification of causal effect assessment and positive and negative fact difference calculation, the scientific nature and accuracy of policy claims data screening are improved. It also provides high-quality data support for the iterative optimization of claims decision models, thereby enhancing the claims decision models' ability to understand claims business logic and improve decision accuracy, thus helping to improve the processing efficiency and rationality of policy claims business.

[0087] Please see Figure 8This application also provides a policy claims data filtering device to implement the above-mentioned policy claims data filtering method. The device includes: The data acquisition module is used to acquire multimodal data of policy claims, which includes policy claim event data and claim result data. The claims prediction module is used to predict claims based on policy claims event data based on a pre-built claims decision model, and obtain model claims decisions. The causal assessment module is used to assess causal effects based on claims outcome data and model-based claims decisions, and to obtain causal effect assessment data. The counterfactual simulation module is used to calculate the difference between positive and negative facts in policy claim event data to obtain the difference data between positive and negative claims. The data filtering module is used to perform high-quality data filtering on multimodal data of insurance policy claims based on causal effect assessment data and positive and negative claims difference data.

[0088] The specific implementation method of this policy claim data filtering device is basically the same as the specific implementation method of the above-mentioned policy claim data filtering method, and will not be described again here.

[0089] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned policy claim data filtering method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0090] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to perform related programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the processing system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and called by the processor 901 to perform the policy claim data filtering method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0091] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described policy claim data filtering method.

[0092] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0093] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0094] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0095] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be placed in one location or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0096] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0097] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0098] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or omitted. The couplings or direct couplings or communication connections shown or discussed may be through some interfaces, or indirect couplings or communication connections between devices or units, and may be electrical, mechanical, or other forms.

[0100] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be placed in one location or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0101] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0103] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for filtering policy claims data, characterized in that, The method includes: Acquire multimodal data on policy claims, wherein the multimodal data on policy claims includes policy claim event data and claim result data; Based on a pre-built claims decision model, claims prediction is performed on the policy claims event data to obtain model claims decisions; Based on the claims result data and the claims decision of the model, a causal effect assessment is performed to obtain causal effect assessment data. The positive and negative factual differences are calculated on the policy claim event data to obtain positive and negative claim difference data; Based on the causal effect assessment data and the positive and negative claims difference data, high-quality data screening is performed on the multimodal data of the policy claims.

2. The method according to claim 1, characterized in that, The causal effect assessment, based on the claims result data and the model claims decision, yields causal effect assessment data, including: The model's claims decision is then subjected to parameter intervention to obtain positive parameter intervention schemes and negative parameter intervention schemes; Based on the aforementioned positive intervention scheme, the expected value of the claims result data is calculated to obtain the expected value of the positive intervention scheme. Based on the aforementioned parameter reverse intervention scheme, the mathematical expectation of the claims result data is calculated to obtain the mathematical expectation data of parameter reverse intervention; Based on the expected mathematical data of positive intervention and the expected mathematical data of negative intervention of the parameters, the causal effect is calculated to obtain the causal effect assessment data.

3. The method according to claim 1, characterized in that, The calculation of the difference between positive and negative facts in the policy claim event data, to obtain the difference data between positive and negative claims, includes: Counterfactual simulations were performed on the policy claim event data to obtain counterfactual claim results; Based on the counterfactual claim results and the claim result data, the difference in claim results is calculated to obtain the difference data between the positive and negative claims.

4. The method according to claim 3, characterized in that, The counterfactual simulation of the policy claim event data to obtain counterfactual claim results includes: Decision variables are extracted from the multimodal data of the policy claims to obtain the claims decision; Based on the aforementioned claims decision, construct counterfactual input samples; Based on the counterfactual input sample, a claims simulation is performed to obtain the counterfactual claims result.

5. The method according to claim 3, characterized in that, The step of calculating the difference in claims results based on the counterfactual claims result and the claims result data to obtain the difference data between the positive and negative claims includes: Feature extraction is performed on the counterfactual claim results to obtain counterfactual claim features; Feature extraction is performed on the claims result data to obtain the actual claims features; The difference between the counterfactual claim features and the factual claim features is calculated using the same dimension feature difference to obtain the positive and negative claim difference data.

6. The method according to any one of claims 1-5, characterized in that, The process of performing high-quality data screening on the multimodal data of policy claims based on the causal effect assessment data and the positive and negative claims difference data includes: Based on the causal effect assessment data, a consistency assessment is performed on the model's claims decision and the claims result data to obtain a data consistency score. Based on the data consistency score, the multimodal data of the policy claims is initially screened to obtain candidate policy claims data; Based on the positive and negative claims data, the candidate policy claims data are screened twice to obtain high-quality policy claims data.

7. The method according to any one of claims 1-5, characterized in that, The pre-built claims decision model predicts claims based on the policy claims event data to obtain model claims decisions, including: Multimodal feature extraction is performed on the policy claim event data to obtain the multimodal features of the policy claim events; The multimodal features of the policy claim event are fused to obtain the standard features of the policy claim event; Based on the claims decision model, feature mapping is performed on the standard features of the policy claims event to obtain the model claims decision.

8. A policy claims data filtering device, characterized in that, The device includes: The data acquisition module is used to acquire multimodal data of policy claims, wherein the multimodal data of policy claims includes policy claim event data and claim result data; The claims prediction module is used to predict claims based on the pre-built claims decision model and obtain model claims decisions; The causal assessment module is used to assess the causal effect based on the claims result data and the model claims decision, and to obtain causal effect assessment data. The counterfactual simulation module is used to calculate the difference between positive and negative facts in the policy claim event data to obtain the difference data between positive and negative claims. The data filtering module is used to perform high-quality data filtering on the multimodal data of the policy claims based on the causal effect assessment data and the positive and negative claims difference data.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the policy claim data filtering method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the policy claims data filtering method as described in any one of claims 1 to 7.