Distributed clustering model training device based on insurance big data and application

By designing a device that combines distributed clustered model training, low-cost cloud computing and large language models, the efficiency and accuracy of insurance big data processing are solved, and efficient and intelligent insurance data processing and model training are achieved.

CN120162671AInactive Publication Date: 2025-06-17王康晟 +2

Patent Information

Application Number
CN202510239574.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle the massive, real-time and accuracy requirements of insurance big data, and traditional model training devices fail to fully consider the complexity and diversity of insurance industry data.

Method used

A distributed clustered model training device based on insurance big data was designed, combining low-cost cloud computing architecture, industry rule guarantee, large language model assistance and human expert feedback mechanism to improve data processing efficiency and intelligent level of model training.

Benefits of technology

It realizes efficient processing of multi-dimensional and highly heterogeneous insurance data, improves the prediction accuracy of the model and the intelligence level of business decisions, and reduces operational costs and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162671A_ABST
    Figure CN120162671A_ABST
Patent Text Reader

Abstract

According to the distributed clustering model training device based on the insurance big data and the application thereof, the distributed clustering model training device based on the insurance big data and the application thereof are combined with distributed computing, cloud computing optimization and a big language model auxiliary strategy, and efficient data processing, low-cost computing resource scheduling and intelligent modeling are achieved. The device adopts a security mechanism to ensure data compliance, introduces man-machine collaborative optimization, improves model suitability and prediction precision, and meets the requirements of the insurance industry for large-scale data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly relates to a distributed clustered model training device and application based on insurance big data. Background Art

[0002] With the rapid development of big data technology, the insurance industry faces unprecedented opportunities and challenges in data processing, analysis, and risk assessment. Insurance companies need to process a large amount of customer data, insurance information, claim records, and other data types, which are widely distributed and complex in terms of time and space. The traditional insurance data processing methods can no longer meet the requirements of efficient storage and real-time analysis of large-scale data. Therefore, how to use advanced computing technology to efficiently process and analyze insurance big data has become an urgent problem in the current insurance industry.

[0003] In recent years, distributed computing and clustered storage technology have gradually become important means to solve the bottleneck of big data processing. Distributed computing can effectively improve the data processing speed and system scalability through the collaborative work of multiple computing nodes. Clustered storage, on the other hand, can provide a larger-scale data storage capacity to ensure that insurance big data can be efficiently accessed and analyzed in real time.

[0004] In addition, with the wide application of machine learning and artificial intelligence technologies, intelligent modeling and prediction based on big data have become the core means to enhance the competitiveness of the insurance industry. Through distributed clustered model training, large-scale data sets can be fully utilized for in-depth learning and training, thereby improving the prediction accuracy of the model and the intelligence level of business decision-making.

[0005] Although some technical solutions have made certain progress in distributed data processing and insurance data analysis, most of the existing technologies have problems such as insufficient computing efficiency, poor system stability, and insufficient data security. Therefore, how to design a more efficient, stable, and powerful distributed clustered model training device has become the key to promoting the intelligent development of insurance industry data.

[0006] Technical problems solved by this patent:

[0007] Lack of model training device designed for the characteristics of insurance big data: The insurance industry involves multi-dimensional and highly dispersed data such as a large amount of customer information, claim records, insurance history, market trends, etc. This data not only has temporality and spatiality, but also often has a high degree of heterogeneity, covering structured data, semi-structured data and unstructured data. Traditional model training devices are mostly based on standardized data processing processes and do not take into account the complexity and diversity of insurance industry data. Therefore, existing training devices are difficult to meet the requirements of the massive volume, real-time nature and accuracy of insurance big data. Through targeting the unique characteristics of big data in the insurance industry, the present invention designs a training device with high-efficient data processing capabilities, which can support the precise training and modeling of multi-dimensional and highly heterogeneous insurance data.

[0008] How to design a cloud computing platform architecture under limited budget: Insurance companies usually do not take AI R & D as their core business, so their cloud computing platform budget is relatively limited. Existing cloud computing platform architectures generally have problems such as high hardware costs and serious resource waste. In order to reduce the platform construction and operation costs, the present invention proposes a unique low-cost cloud computing architecture design. By optimizing the resource scheduling and computing task allocation strategies, the utilization efficiency of computing resources is maximized, and a distributed resource pool is used for flexible scheduling, so as to achieve the maximum utilization of the limited budget of insurance companies. At the same time, the cloud computing architecture of the present invention can be dynamically expanded, and resources can be elastically scheduled according to specific computing requirements to ensure that while ensuring computing performance, the operation costs are controlled.

[0009] Industry-specific rules and data security control issues in the insurance industry: When dealing with customer data and business data, the insurance industry often involves strict confidentiality regulations and privacy protection requirements. Insurance companies need to ensure that they meet industry compliance and confidentiality requirements in all aspects of data use, storage and sharing. When designing the model training device, the present invention fully considers the unique rules of the insurance industry, proposes a secure distributed architecture, and guarantees data security through means such as data encryption, access control, and data desensitization. At the same time, while realizing data processing and model training, this device embeds a security mechanism that meets the regulatory requirements of the insurance industry to ensure that industry regulations and confidentiality agreements are strictly implemented.

[0010] Introduction of large language model assisted strategy: In the traditional model training process, specific rules are usually set manually or rely on static training data, failing to effectively combine real-time feedback and intelligent decision-making. Through the embedding of large language models, this invention provides strategy assistance support throughout the production chain of model training. In each link such as data preprocessing, model training, and prediction evaluation, the natural language processing ability of large language models is utilized to provide semantic-level understanding, data cleaning suggestions, and business logic optimization for the model training process, thereby improving the intelligence level of the training process and further enhancing the accuracy and reliability of business models in the insurance industry.

[0011] Introduction of human expert feedback mechanism: This device innovatively introduces a human expert feedback mechanism to perform real-time intervention and adjustment on the model throughout the model training process. By embedding the input and output of physical devices during the training process, experts can provide real-time feedback at various stages of the model, ensuring that problems during the training process can be discovered and corrected in a timely manner. The feedback from experts not only helps optimize model parameters but also provides more accurate guidance for the model from the perspective of industry experience, enhancing the practicality and business adaptability of the model. The real-time and interactive nature of this feedback mechanism effectively improves the flexibility and accuracy of model training. Especially when facing complex industry data, it can quickly adjust strategies to adapt to the dynamic changes of the industry. Summary of the Invention

[0012] Aiming at various technical problems in existing insurance big data processing, low-cost cloud computing platform design, industry rule compliance guarantee, application of large language model assisted strategy, and real-time feedback mechanism of human experts, etc., this invention provides a distributed clustered model training device and application based on insurance big data.

[0013] The technical solution of this invention provides a distributed clustered model training device and its application based on insurance big data. This invention relates to a distributed clustered model training device and application based on insurance big data, specifically an intelligent model training system that combines the characteristics of the insurance industry, a low-cost cloud computing architecture, industry rule guarantee, large language model assistance, and a human expert feedback mechanism. The system aims to improve the intelligence level, computing efficiency, and compliance of the insurance industry in big data processing and analysis, and solve the deficiencies of existing technologies in large-scale data processing and business model training.

[0014] The implementation of this invention includes the following parts:

[0015] A. The data causality check module A0 covering the entire production chain audits all data change processes in the entire business process, enhancing the relationship between variables from correlation to causality in the whole process of calculation - prediction - analysis. Specifically, it includes: an optimized insurance data audit method enhanced by a graph network and an optimized dynamic deduction data audit method combining virtual and real. The implementation steps are as follows:

[0016] A1. Considering that the variables in the insurance data calculation process are relatively fixed, and there is a basically stable functional relationship between the implicit physical, practical, and economic meanings among the variables, but the functional relationship is not explicit and is often not in polynomial form, the present invention first uses a graph network to model the relationship between variables to achieve enhanced representation of the relationships among various variables, and uses the constructed graph network to conduct causality evaluation and audit on the data throughout the production process. The specific method is as follows:

[0017] A11. Select reliable and true data that has passed historical verification as Ground - Truth as the initial data for network modeling. The data features include: policyholder ID, gender, age, driving experience, credit score, historical violation records, vehicle brand, vehicle model, vehicle age, displacement, mileage, number of claims in the past five years, repair location, agent ID, whether frequently using the same repair shop, social network relationship, earned premium of auto insurance, claim frequency, and amount of each claim.

[0018] A12. In the process of insurance data processing, due to the highly non - linear, non - explicit nature of the variable relationships involved and the implicit economic, physical, and social rules, a dynamic graph network modeling method based on implicit function relationships is combined with a custom - defined causal inference algorithm to deeply abstract and reason about the causal relationships in insurance big data.

[0019] A121. Dynamic graph network and causal relationship modeling

[0020] First, each variable in the insurance data is regarded as a node in the graph, and the features of the node are embedded through a set of embedding functions f θ (x), where x represents the data points in the insurance business (such as customer personal information, policy terms, historical claim data, etc.). Different from traditional graph convolution operations, the graph structure update method of the present invention does not depend on a fixed adjacency matrix or a static graph structure, but dynamically adjusts the edge weights between nodes based on an implicit causal model.

[0021] Set the feature representation of each node v i as h i , and its update process is described by the following formula:

[0022]

[0023] Among them, represents the set of adjacent nodes of node v i at time t, is the dynamic adjustment coefficient of the connection strength between nodes. Its calculation is based on the implicit temporal relationship in the data, historical data, and the results of the causal inference model; ψ t (h i (t) , h i (t-1) ) is the time-dependent term generated based on the causal inference algorithm, capturing the changing trend of node features. α t is a trainable parameter, and f θ represents multiple alternative function types, including polynomial functions, exponential functions, logarithmic functions, and fractional functions.

[0024] A122. Causal Inference and Adaptive Causal Relationship Function

[0025] To effectively improve the causal inference between nodes in the graph network, the present invention introduces an adaptive causal inference function during the network update process. This function can dynamically learn the causal relationship between variables according to the evolution of the data stream. Specifically, an implicit causal model is designed, and the core idea is to optimize the network parameters by calculating the causal path probability between each pair of nodes.

[0026] In each round of update, the causal inference process of the node is carried out according to the following recursive formula:

[0027]

[0028] Among them, Y represents the target variable (such as the final state of an insurance claim), and X i is the input feature of the i-th node (variable), and C i are the conditional variables introduced during the temporal and causal inference processes, which describe the influencing factors at the current moment. is the conditional probability distribution, representing the occurrence probability of the target variable Y given the historical data X and the causal condition C.

[0029] A123. Causal Path and Conditional Independence Constraint

[0030] To further improve the accuracy of causal inference, the present invention uses causal path constraints to construct a complex dependency network. For each pair of nodes v i and v j , by calculating their causal path probabilities and performing conditional independence constraints, the connection weights of the edges are further optimized. Specifically, the present invention calculates the causal path between nodes through the following formula:

[0031]

[0032] Among them, represents all possible path sets, is a conditional independence constraint used to control the rationality of causal paths. k represents the path between nodes.

[0033] A124. Temporal nesting and collaborative optimization mechanism

[0034] Considering the significant impact of the changes in time series in insurance data on causal inference, the present invention introduces a collaborative optimization mechanism based on temporal nesting. This mechanism performs weighted control in the time dimension by introducing a recursive time gating function G t to dynamically adjust the influence of time series data on causal inference. Specifically, the update function of the node is:

[0035]

[0036] Among them, G t is a time gating function used to control the degree of dependence of node updates on causal inference paths at each time step. β t is a trainable parameter.

[0037] A125. State transition function between nodes and causal inference

[0038] In the graph network model of the present invention, the state transition between nodes is designed as a highly non-linear and dynamically changing process, aiming to perform adaptive learning and optimization according to the complexity and implicit causal relationships of insurance big data. The state transition of nodes does not only depend on the simple linear relationship between nodes, but reflects more complex variable change laws by integrating historical data, causal inference, and time series information.

[0039] First, consider the state h i of each node v i in the insurance data. Its change over time step t can be expressed as a recursive process. The change of node state is affected by the intrinsic characteristics of the node (such as customer information, policy type, historical claim data, etc.) and the causal relationship between nodes. A new state transition function is introduced to describe the causal state transition from node v i to node v j at time t. Use the Markov chain function as the basic function for enhanced training:

[0040]

[0041] Among them, h i (t) and h j (t) are the states of node v respectivelyi and v j The feature representation at time t is conditional information containing the causal relationship between nodes v i and v j This state transition function describes the evolution of the causal relationship between two nodes and can be updated according to historical information and causal paths at different time steps.

[0042] A125.1 State Transition under Implicit Causal Model

[0043] To more precisely capture the state transition in the causal reasoning process, the present invention further optimizes the dynamic update of node states by introducing an implicit causal model. Assume that the state transformation of node v i is not only affected by the input feature h i (t) of the node itself, but also affected by the historical state h i (t-1) and the external environment C i . The state transition function can be expressed as:

[0044] h i (t+1) = h i (t) + f θ (h i (t-1) , C i )

[0045] where f θ is a non-linear mapping function used to represent the change relationship of the state of node v i between time t and time t-1. By embedding the feature information of multiple historical moments, this function can adapt to complex time series data and potential non-linear causal relationships.

[0046] A125.2 Structured Causal Path Constraint and Optimization

[0047] To further improve the reasoning accuracy, for the causal state transition between nodes, the present invention introduces a structured causal path constraint. This constraint not only considers the direct causal relationship between nodes, but also includes the high-order dependence paths of nodes. For this purpose, a causal path function is defined to describe the causal path from node v i to node v j :

[0048]

[0049] where γ ijk is the coefficient in the causal path, which represents node v iTo node v j The high - order influence relationship, where high - order refers to the influence relationship formed based on high - order inference logic as opposed to first - order inference logic.

[0050] A125.3 Timing gating mechanism and causal path adjustment

[0051] Considering the important influence of timing factors in insurance data on causal inference, the present invention introduces a timing gating mechanism during the state transition process. This mechanism can precisely control the timing characteristics of information flow by dynamically adjusting the causal path weights between different time moments. The timing gating mechanism is implemented through the following function:

[0052]

[0053] where G t is a gating function used to adjust the strength of the information flow on the causal path at time step t. By adjusting the value of G t , the model can adaptively respond to the dynamic changes in the time series and map this response to the result of causal inference.

[0054] A125.4 Causal inference and final prediction

[0055] Through the above - mentioned node state transition and causal path optimization mechanism, each node v i in the network can be updated at each moment based on historical data, causal paths, and timing information, and reflect the evolution process of each node in the entire system. Finally, by comprehensively reasoning about the states of all nodes, the prediction results of each decision point in the insurance data can be obtained. The output of this process is the reasoning result of business problems such as risk assessment, claim prediction, and customer behavior analysis in the insurance industry.

[0056] The final prediction result can be expressed by the following formula:

[0057]

[0058] where H i represents the final state of node v i , and P(Y i |H i ) is the prediction probability based on the node state. T represents the time step used for prediction reference.

[0059] A125.5 It is accelerated by hardware of a specifically designed insurance data audit optimization method dedicated to graph network enhancement.

[0060] Due to the excessive use of symbols and limitations, the mathematical symbols used in A12 are only represented in A12 and have no relation to the symbols with the same representation described in other steps.

[0061] A13. In addition to auditing the data throughout the entire process of calculation - prediction - analysis by means of graph modeling, the present invention also provides an optimization method for dynamic deductive data auditing that combines virtual and real elements. As a coordination, this method conducts virtual simulation modeling on actual data. Using data such as the water leakage situation of the house roof, the construction years of houses in the block, and the historical risk occurrence situations in the same block as references, it performs dynamic deductive reasoning and applies auditing to the data by means of counterfactual reasoning, specifically as follows:

[0062] A131. First, through virtual simulation modeling of actual data, a dynamic deductive model corresponding to real - world data is constructed. Factors such as the water leakage situation of the house roof, the construction years of houses in the block, and the historical risk occurrence situations are set as input data. These data will serve as the initial conditions of the virtual simulation model, and through deductive reasoning, the evolution process of the house under different time nodes and different conditions is generated.

[0063] Specifically, the construction of the virtual simulation model can be expressed as:

[0064] X t =f(X t-1 ,C t ,E t )

[0065] where X t represents the virtual data state at time step t, C t is the external condition at time t (such as the block environment, house age, etc.), and E t is the error term or random perturbation in the deductive reasoning process. The function f represents the dynamic deductive process, which is based on the theory of contradiction and Socrates' syllogism. It describes the law of the change of the virtual data state over time and is fitted using a periodic function.

[0066] A132. On the basis of virtual simulation modeling, the present invention further introduces the counterfactual reasoning method to audit the data. The core idea of counterfactual reasoning is to infer "how the result would change if the values of certain variables were changed" by simulating hypothetical situations. This method is particularly applicable in the insurance field because it can, in the absence of a clear causal chain, evaluate the possible results in different scenarios by simulating different causal paths and avoid Simpson's paradox.

[0067] Set the counterfactual reasoning model as:

[0068]

[0069] where, represents the counterfactual prediction result of node i, f cfIt is a counterfactual reasoning model that generates conditions and environments through virtual simulation modeling and dynamic deductive reasoning to obtain results under hypothetical scenarios.

[0070] Using counterfactual reasoning, it is possible to simulate data changes under different scenarios. For example, by changing the likelihood of a house roof leaking, it can simulate whether a house will have a leaking problem under different construction years, thereby auditing potential risks. This counterfactual reasoning method helps identify potential problems in the data and optimize and adjust the prediction results.

[0071] A133. Optimization of data auditing The optimization process of data auditing consists of the following steps:

[0072] A133.1 Data collection and preprocessing: Collect real-world observational data containing multi-dimensional data and standardize and clean it to ensure the reliability and integrity of data quality.

[0073] A133.2 Virtual simulation modeling: Based on the collected actual data, construct a virtual simulation model. This model maps the actual data through dynamic simulation and virtual environment to support further reasoning and analysis of the data.

[0074] A133.3 Counterfactual reasoning: Use counterfactual reasoning to deduce virtual data and speculate on how the data would change if certain conditions changed. Specifically, the process of counterfactual reasoning can be represented by the following general mathematical function:

[0075] Let X i be the actual observational data of data point i, which contains multiple variables (such as the construction year of the house, environmental impact, etc.), C i be the external environmental factors (such as regional policies, historical risks, etc.), and A i be the hypothetical condition (such as changing the value of a certain variable), then the counterfactual reasoning model can be described by the following function:

[0076]

[0077] where, represents the counterfactual prediction result of data point i under the given virtual hypothesis condition, and f cf is the counterfactual reasoning function, which generates new predicted values by combining the actual observational data X i , the external environment C i , the error term E i and the hypothetical change factor A i .

[0078] For example, assume that X i contains the historical performance data of a certain system, and A iis a change in an operating parameter of the system (such as changing the parameter value A param ), and the impact of this parameter change on system performance (such as failure rate, efficiency, etc.) can be deduced. At this time, the counterfactual reasoning process is described as:

[0079]

[0080] By changing the hypothetical condition A param , different counterfactual results can be generated, thereby evaluating the system performance under different assumptions.

[0081] A133.4, Auditing and Optimization: Based on the results of counterfactual reasoning, conduct data auditing to identify potential errors, defects, and inconsistencies in the data. Through the audit results, further optimize the data model and reasoning algorithm to improve the accuracy and consistency of the prediction results.

[0082] A134, Mapping from Virtual Environment to Real Business Scenarios

[0083] A134.1 Virtual Data Generation: In the previous steps, simulation data has been generated through a virtual simulation model This data is obtained through the deduction of simulating real data in a virtual environment, and the simulation data can be expressed as:

[0084]

[0085] where g sim is a simulation generation function that generates output data in a virtual environment based on the input real data X i , the external environment C i and the hypothetical condition A i .

[0086] A134.2, Data Mapping Optimization Goal: To achieve the mapping of simulation data to real-world data, a mapping function M sim2real needs to be constructed. Through this function, the simulation data is converted into real-world data This function can be expressed as:

[0087]

[0088] where B i is an additional correction parameter in actual observations, representing the conversion error or deviation from the simulation model to the real world, and C i represents the external environment. This mapping function makes the simulation data consistent with real-world data in terms of physical quantities, spatio-temporal relationships, and other relevant data characteristics by reverse optimization and correction of the correction factor.

[0089] A134.3, Optimization Algorithm and Mapping Precision Enhancement: To improve the mapping precision, real-world data is used to optimize the mapping function M sim2real .

[0090] The mapping precision is optimized by minimizing the following objective function:

[0091]

[0092] The minimization process of this objective function can be achieved through the gradient descent method, and the parameters of the mapping function are adjusted by backpropagating the error.

[0093] A134.5, Reverse Correction and Amendment: In the process of converting simulation data into real-world data, the influence of the external environment and other uncertain factors need to be considered. By introducing a reverse correction mechanism, the mapping precision is further optimized. For example, historical data or expert knowledge can be used to adjust the parameter B in the mapping function i , making the model more adaptable to the complex factors in the real-world environment.

[0094] A134.6, Final Data Audit and Optimization: Finally, based on the optimized data mapping function M sim2real , the simulation data and the real-world data are accurately docked, and through the data audit process, the data consistency and validity are checked, potential anomalies or errors are identified, thereby optimizing the data quality and improving the accuracy of prediction and decision-making.

[0095] Due to the excessive and limited use of symbols, the mathematical symbols used in A13 are only represented in A13 and have no relation to the symbols with the same representation described in other steps.

[0096] B. A Hardware Acceleration Method B0 for an Insurance Data Audit Optimization Method Dedicated to Graph Network Enhancement, specifically including: a hardware accelerator architecture B-inf dedicated to insurance data audit for graph network enhancement and its acceleration process. This hardware accelerator realizes efficient graph data processing, storage, and optimization calculation through a customized multi-core processor, a dedicated graphics processing unit (GPU), and large-scale parallel computing capabilities, fully coupling the algorithm and the hardware.

[0097] Hardware Components:

[0098] B1. Dedicated Processing Unit (PU, Processing Unit): This hardware accelerator design includes multiple dedicated processing units, which are respectively used for the creation and update of graph data structures, the enhanced calculation of graph networks, and the in-depth audit optimization of graph data. Each PU can parallelly process different graph node and edge information, supporting operations such as efficient graph traversal, shortest path calculation, and connectivity analysis.

[0099] B2. Graphics Processing Unit (GPU): Through the integrated Graphics Processing Unit (GPU), the hardware can efficiently execute the training and inference calculations of parallel graph networks. By implementing parallel computing at the hardware level, the GPU supports the graph modeling and optimization of large-scale insurance data, significantly reducing the bottleneck of traditional CPU computing.

[0100] B3. Graph Storage Unit (GSU): To support the storage and reading of large-scale insurance data graphs, the present invention designs a dedicated Graph Storage Unit (GSU). This unit adopts a high-bandwidth storage architecture, supports the storage of large-capacity graph-structured data, and can perform fast data exchange between the memory and the processing unit, ensuring the efficient access of graph data.

[0101] B4. Hardware Acceleration Interface (HAI): This hardware accelerator designs a dedicated Hardware Acceleration Interface for data transfer and interaction with external systems or data sources. Through the HAI interface, insurance data can be quickly transferred to the hardware accelerator for processing, and at the same time, the processing results can be returned in real time to support real-time data auditing.

[0102] Hardware acceleration process:

[0103] B5. Insurance data input: External insurance data is input into the hardware accelerator through the HAI interface. The input data includes historical insurance claim records, customer information, insured product data, etc.

[0104] B6. Graph network construction: The dedicated PU will receive the input insurance data and construct an insurance data graph, which contains nodes (such as customers, policies, claim events) and edges (such as association relationships, claim histories, etc.). This process is executed in parallel by the graphics processing unit in the hardware accelerator and can quickly complete the construction of large-scale data graphs.

[0105] B7. Graph enhancement and optimization: After the graph data is constructed, the GPU performs deep learning training and auditing optimization on the graph based on module A0. The optimization tasks include predicting potential risks, analyzing customer behavior, identifying possible abnormal events, etc.

[0106] B8. Output audit results: The processed audit results are transmitted back to the external system through the HAI interface to ensure the real-time and accuracy of the audit results.

[0107] C. A hardware acceleration method C0 for a dynamic deduction data audit optimization method dedicated to the combination of virtual and real, specifically including: a hardware accelerator architecture C-inf that integrates virtual simulation and physical data processing and its acceleration process, which comprehensively supports the interaction and optimization of dynamic simulation and real-world data by combining high-performance computing units, physical simulation acceleration modules, and dynamic mapping processing units.

[0108] Hardware components:

[0109] C1. Simulation Processing Unit (SPU): The simulation processing unit is specifically used for the generation and processing of virtual simulation data. Driven by the real-world business data simulation model, this unit can generate simulation data based on the input real-world data, support dynamic parameter adjustment and environmental simulation. The SPU can simulate different simulation scenarios according to real-time input conditions.

[0110] C2. Data Mapping Accelerator (DMA): The data mapping accelerator is responsible for the mapping between simulation data and real-world data. Through a customized mapping algorithm (such as reverse optimization and correction mechanism), this accelerator realizes the efficient conversion from virtual data to real-world data. The DMA accelerator supports large-scale data mapping and can quickly respond to real-time data input.

[0111] C3. Dynamic Fusion Module (DFM): This module is used to process the fusion of virtual and real-world data. The DFM can dynamically combine virtual simulation data with real-world data, automatically adjust simulation parameters to adapt to changes in the real-world environment, and improve the accuracy of data fusion through optimization algorithms.

[0112] C4. Multi-dimensional Processing Compute Unit (MPCU): This computing unit is designed to efficiently execute complex multi-dimensional dynamic deduction computing tasks. By combining a hardware-level parallel computing architecture, the MPCU can process multiple types of data simultaneously, supporting the parallel processing of spatio-temporal data, dynamic deduction inference results, and counterfactual inferences.

[0113] C5. Hardware Collaborative Interface (HCI): This interface is used for data exchange with external hardware or cloud computing platforms. The HCI interface supports high-speed data transmission, ensuring that data can be seamlessly transferred from the hardware acceleration platform to external systems for further analysis and decision support.

[0114] Hardware acceleration process:

[0115] C6. Real - world data input: External real - world data is input into the hardware platform through the HCI interface. The data content includes real - time monitoring data, environmental data, etc.

[0116] C7. Virtual simulation generation: The simulation processing unit (SPU) generates simulation data in the virtual environment based on the input real - world data. The simulation model generates an expected data scenario based on the physical model, historical data, and external variables.

[0117] C8. Mapping of simulation data and real - world data: Through the data mapping accelerator (DMA), the system docks the simulation data and real - world data to ensure that the simulation results can accurately reflect the dynamic changes in the real world.

[0118] C9. Dynamic deduction and optimization: The dynamic data fusion module (DFM) performs dynamic deductive reasoning on the data through backward reasoning and correction mechanisms. The optimization process includes compensation for uncertain factors and adjustment of the deductive model.

[0119] C10. Output of audit results: The final audit results are output to the external system through the HCI interface for generating decision - support reports or for real - time monitoring.

[0120] D. An audit module D0 built on the basis of the architectures B0 and C0 described in methods B0 and C0, specifically including:

[0121] D1. Processing Unit (PU): This unit is responsible for performing the computational tasks of graph network enhancement and dynamic deductive data audit with the combination of virtual and real data. The processing unit includes multiple dedicated computing cores for parallel processing of complex graph data structures and simulation calculations to ensure efficient data processing and reasoning optimization during the audit process.

[0122] D2. Memory Unit: The memory is used to store intermediate calculation results related to graph data, historical audit data, virtual simulation models, and the results of counterfactual reasoning. Through high - capacity and high - bandwidth storage, it supports high - speed read and write operations to ensure no bottlenecks occur during the execution of complex computational tasks.

[0123] D3. Simulation Module: This module is specifically used to generate and process virtual simulation data according to the virtual - real combination scheme in method C. The simulation module can, with the support of hardware acceleration, generate virtual scenarios based on real - world data in real time to simulate the audit process in different environments.

[0124] D4, Data Mapping Accelerator (DMA): This component is responsible for performing mapping calculations between virtual simulation data and real-world data, ensuring seamless data conversion from the simulation environment to the real-world environment, and guaranteeing the authenticity and accuracy of the final audit results.

[0125] D5, Interface Module: Used for data interaction between the audit server and external systems or network environments. The interface module supports high-speed data transmission and instruction execution, ensuring that the input process data of the entire business chain can be quickly transmitted to the audit server and promptly returned to the external system after the audit results are generated.

[0126] D6, Graph Data Processing Unit: This unit is specifically used to process graph network-related tasks, including graph data construction, node and edge processing, graph enhancement and optimization, etc. Through hardware-level parallel computing, the graph data processing unit can significantly improve the efficiency and accuracy of graph data operations.

[0127] D7, Dynamic Inference Module: The dynamic inference module is responsible for performing counterfactual reasoning and dynamic deductive analysis according to the graph network and dynamic deductive algorithm in Method B. By combining the spatio-temporal data model, it can optimize the data audit work in the entire business chain process.

[0128] The memory of this module stores the computer program D1. When the computer program D1 is executed, the module audits the process data of the entire business chain according to Method A. The computer program D8 contains an instruction set for coordinating the collaborative work of the above-mentioned hardware components to achieve a full-process audit of data input, processing, simulation, optimization, and output. Specifically, when the program D8 is executed, the audit server first receives the external business chain process data and transmits it to each processing unit through the data interface module. Then, the graph data processing unit constructs and optimizes the graph network model, the simulation module generates a virtual scenario based on real-time data, and realizes the data conversion of the combination of virtual and real through the data mapping accelerator. Finally, the dynamic inference module performs in-depth audit analysis on the data through counterfactual reasoning and outputs an audit report.

[0129] E. A distributed storage method E0 for insurance data after graph modeling is specifically implemented as follows:

[0130] This method mainly realizes distributed storage and computing integration of data by fragmenting and storing the insurance data after graph modeling and tightly coupling the storage process with the computing process. The specific steps are as follows:

[0131] E1. Data Sharding: When performing graph modeling, first shard the insurance data according to certain partitioning rules. Each shard of data contains node and edge information within a certain range, ensuring uniform distribution of data across different computing nodes. Assume the graph data is where is the set of nodes and E is the set of edges. The data sharding process can be expressed as:

[0132]

[0133] where is the i-th data shard, N is the total number of shards, and each contains the node set and the edge set E i . The sharding rules can be determined based on node degrees, the topological structure of the graph, or edge weights.

[0134] E2. Node-to-Compute Mapping: Store the sharded data on different computing nodes to ensure that each computing node can store the corresponding data fragment and can execute graph data computing tasks locally. Assume there are M computing nodes Each computing node stores the shard Then the mapping relationship can be expressed as:

[0135]

[0136] This mapping is dynamically adjusted according to load balancing and data dependencies to ensure that the computing tasks processed by each computing node minimize cross-node communication.

[0137] E3. Compute-Storage Co-location: On each computing node, not only the sharded data is stored, but also the associated computing units are included. The computing task of each computing node can be expressed as:

[0138]

[0139] where f(v) represents the computation performed on node (such as node feature extraction), and g(e) represents the computation performed on edge e ∈ E i (such as edge weight calculation). The computation for each node is limited to the locally stored data, ensuring that the computation process does not involve cross-node data access, thereby improving computing efficiency.

[0140] E4. Data Synchronization & Consistency: Since data is stored on different computing nodes, this method introduces an efficient distributed synchronization mechanism to ensure data consistency among all computing nodes. The data synchronization process is achieved through a consistency protocol, such as the Paxos or Raft protocol. Suppose data is modified on node The synchronization process can then be expressed as: where

[0141]

[0142] is the data on node is the increment of the modification, and is the data merge operation. After synchronization, the data on all computing nodes will be updated to a consistent state. is the increment of the modification, is the data merge operation. After synchronization, the data on all computing nodes will be updated to a consistent state.

[0143] E5. Data Query and Distributed Query Engine: To support efficient data query in a distributed storage environment, this method designs a distributed query engine. The query request of the query engine can be expressed as:

[0144]

[0145] where Q is the query operation, is the graph data, are the query parameters (such as path query, node feature query, etc.), and is the query result. During the query process, the query engine will automatically determine the storage location of the query data and distribute the request to the corresponding computing nodes for processing. The merging of cross-node queries is done through:

[0146]

[0147] where is the query result returned by computing node and is the final query result.

[0148] E6. Fault Tolerance: In a distributed storage system, to ensure that the system can still work properly in case of failures or outages of some computing nodes, this method introduces a fault tolerance mechanism. Data backup can be achieved through a replication strategy. Suppose there are K replicas of data on multiple computing nodes, which can be expressed as:

[0149]

[0150] Among them, is the k-th copy on the node If a certain node fails, the system will automatically switch to the data on this copy for calculation to ensure that the task can continue to execute.

[0151] This method can realize the integration of distributed storage and calculation, ensure the efficient storage, fast calculation and query of large-scale insurance data, and at the same time maintain the consistency and reliability of data in a dynamic environment. Through this method, the storage and calculation process of insurance data after graph modeling no longer depends on centralized computing resources, can effectively share the computing load, and reduces the pressure on a single computing node. This method is particularly suitable for the insurance industry that needs to process massive data and has complex computing requirements, and can improve the speed and accuracy of data processing.

[0152] Due to the excessive and limited use of symbols, the mathematical symbols used in E are only represented in E and have nothing to do with the symbols with the same representation described in other steps.

[0153] F. An insurance data processing method F0 based on the Generalized Linear Model (GLM)

[0154] After data graph modeling, in the data processing stage, aiming at the characteristics of insurance data, method F is designed based on GLM, which specifically includes:

[0155] F1. Method overview

[0156] This method is based on the Generalized Linear Model (GLM, Generalized Linear Model), combined with the business characteristics of the insurance industry, to achieve accurate prediction and optimized processing of insurance data. This method covers multiple links such as data preprocessing, feature engineering, model training, and prediction optimization, and is applicable to multiple insurance business scenarios such as auto insurance pricing, health insurance risk assessment, life insurance actuarial, and fraud detection.

[0157] F2. Data preprocessing

[0158] In the Generalized Linear Model (GLM) method, data preprocessing is crucial because it directly affects the stability and prediction accuracy of the model. Aiming at the characteristics of insurance industry data, the preprocessing link of this method includes the following steps:

[0159] F21. Data cleaning The purpose of data cleaning is to ensure the integrity, correctness and consistency of data, and avoid interference from outliers, missing values and incorrect data to model training. The specific cleaning measures include:

[0160] F21.1. Removing duplicate data

[0161] Insurance data may come from multiple systems, such as the insurance application system, claims settlement system, etc., and there are duplicate records. The removal methods include detecting fields such as the same applicant ID, license plate number, insurance policy number, etc. If all field contents are the same, the duplicate items are deleted.

[0162] F21.2, Handling missing values

[0163] For continuous variables (such as age, credit score), fill them with the mean, median, or mode.

[0164] For categorical variables (such as car brand, agent ID), fill them with the mode (the most common category) or set them to the "unknown" category.

[0165] F21.3 Detect and handle outliers Calculate the interquartile range (IQR) and remove values outside the range of [Q1 - 1.5IQR, Q3 + 1.5IQR].

[0166] If the data follows a normal distribution, remove values exceeding the mean ± 3 times the standard deviation.

[0167] Use Winsorization to truncate outliers and bring them back to a reasonable range.

[0168] F22, Feature selection The purpose of feature selection is to eliminate redundant or irrelevant variables to reduce computational complexity and improve the generalization ability of the model. The main methods include:

[0169] F22.1, Correlation analysis

[0170] Calculate the correlation between feature variables and the target variable, and select variables with stronger correlations. For example: If it is found that the vehicle color has no obvious impact on the claim amount, then this variable can be deleted.

[0171] F22.2, Variance filtering

[0172] Calculate the variance of each feature variable and eliminate variables with too low variance (i.e., almost all sample values are the same).

[0173] F22.3, Recursive feature elimination (RFE)

[0174] Train a preliminary GLM model, calculate the regression coefficient (β) of each variable, and gradually remove features with less contribution to the prediction.

[0175] F22.4, Combining business expert knowledge

[0176] In some cases, relying solely on statistical methods to select features may be insufficient. It is necessary to combine industry experience and use an intelligent insurance data processing and decision optimization method H0 based on large language models and human expert feedback.

[0177] F22.5. The age of the policyholder affects the claim rate.

[0178] Past claim records are crucial for fraud detection. Social network relationships reflect potential group fraud behavior.

[0179] F23. The GLM with variable transformation requires the input data to meet specific assumptions (such as normal distribution, linear relationship, etc.), so some variables need to be transformed:

[0180] F23.1. Categorical variable encoding

[0181] · One - Hot Encoding: Suitable for unordered categories, such as "car brand" (Toyota, Honda, BMW).

[0182] · Label Encoding: Suitable for ordered categories, such as "credit rating" (A, B, C).

[0183] · Target Encoding: Calculate the target mean for each category, for example, the impact of "insured area" on the claim rate.

[0184] F23.2. Numerical variable transformation

[0185] · Log Transformation: Used for right - skewed distribution data. For variable x, such as claim amount:

[0186] x ′ = log(1 + x) (1)

[0187] · Square Root Transformation, for variable x, suitable for slightly right - skewed data:

[0188]

[0189] · Box - Cox transformation of x, used to make the data closer to normal distribution and improve the model stability.

[0190] F23.3. Standardization

[0191] Standardization is used to eliminate the difference in variable scales and avoid the dominant role of some variables in the loss function during model training. The following standardization methods are adopted for variable x:

[0192] · Z - score standardization

[0193]

[0194] Among them, μ is the mean and σ is the standard deviation.

[0195] · Min - Max normalization Applicable to variables within a specific range, such as credit scores (usually between 300 - 900).

[0196] · Robust Scaling Among them, Q1 and Q3 are the first and third quartiles respectively, applicable to variables with extreme values, such as claim rates.

[0197] F23.4. Handling data imbalance

[0198] In insurance claim rate prediction, data may have imbalance problems (e.g., most customers have 0 claims, and only a few customers make claims). The following are optional methods:

[0199] · Undersampling

[0200] Randomly remove some majority - class samples to make the ratio with the minority class more balanced. The undersampling method randomly removes majority - class samples, reducing the number of majority - class samples, thereby achieving class balance in the data. For example, in a dataset, 95% of the customers have no claims, and only 5% of the customers make claims. Through undersampling, some customer samples with no claims can be randomly selected and removed to make the number of the two types of samples closer.

[0201] · Oversampling

[0202] The oversampling method balances the dataset by synthesizing minority - class samples. Using the SMOTE (Synthetic Minority Over - sampling Technique) method, new synthetic samples can be generated between minority - class samples through interpolation techniques. For example, assume that in a dataset, only 5% of the customers make claims. Through the SMOTE method, new minority - class samples can be generated to enhance the model's ability to identify the minority class.

[0203] · Weighted loss function

[0204] During the GLM training process, higher weights are assigned to minority classes to reduce the impact of class imbalance on the model. During the training process, a weighted loss function is used to assign higher weights to minority classes, reducing the impact of class imbalance on the model. This method can make the model more sensitive to minority class data during the training phase and improve its classification accuracy. For example, in a Generalized Linear Model (GLM), the weights of customers who file claims can be increased, making the model pay more attention to these less frequent customer samples during training.

[0205] In practical applications, appropriate methods can be selected according to the characteristics of the data. For example, if the minority class samples in the dataset are too sparse, the SMOTE oversampling method can be used to synthesize more samples and improve the model's prediction ability for minority classes. In the case of relatively balanced sample numbers, using a weighted loss function may be more appropriate because it can improve the model's performance without increasing the sample size.

[0206] F23.5. Data PartitioningData partitioning is used to ensure the generalization ability of the model, and it consists of three parts: the training set, the validation set, and the test set:

[0207] · Training Set (Train Set, 70%): Used for model training.

[0208] · Validation Set (Validation Set, 15%): Used to adjust hyperparameters.

[0209] · Test Set (Test Set, 15%): Used for final evaluation.

[0210] Partitioning Method:

[0211] · Out-of-time Test Set: Use the dataset in the future of the training set for testing.

[0212] F3. Selecting the GLM ModelSelect a suitable GLM variant according to the characteristics of the target variable:

[0213] · Poisson Regression: Suitable for predicting claim frequency, i.e., the number of claim occurrences. The Poisson regression model can effectively describe the distribution of the number of event occurrences and is often used to analyze scenarios where the number of claims is small but occurs frequently. For example, use the Poisson regression model to predict the number of claims made by insurance customers within a certain period.

[0214] · Gamma Regression: Suitable for predicting claim amount, i.e., the average payout cost per claim. Gamma regression is suitable for modeling continuous data with a positive skewed distribution, such as claim amounts which are usually positive and have large variability. For example, the Gamma regression model can predict the average compensation cost for a claim.

[0215] · Log-Normal Regression: Suitable for predicting claim amounts with large fluctuations. The log-normal regression model can handle claim amounts with large fluctuations or long-tailed distributions and is suitable for situations where the claim amounts are high and the distribution is asymmetric. For example, predicting the likelihood of high-value payouts and the compensation amounts.

[0216] · Tweedie Regression: Suitable for simultaneously modeling mixed data of zero claims and positive payout amounts. Tweedie regression can handle situations with a large number of zero values (no claims occurred) and a small number of positive values (claims occurred) and is very suitable for the zero-inflation problem in insurance claim prediction. For example, using the Tweedie regression model to predict the probability of a customer making a claim and the compensation amount.

[0217] The above are optional variants. In this method, the present invention adopts Poisson regression and Gamma regression to predict the claim frequency and claim amount respectively. In this method, the present invention uses two models, Poisson regression and Gamma regression, which are respectively used to predict the claim frequency and claim amount. Poisson regression is used to model the number of claims, while Gamma regression is used to predict the average payout amount per claim.

[0218] F4. Model Training This method uses the Maximum Likelihood Estimation (MLE) to solve for the parameters, and the model training process is as follows:

[0219] F4.1. Define the GLM model

[0220]

[0221] Among them, g(·) is the link function, X is the input variable, Y is the target variable, and β is the parameter to be estimated.

[0222] F4.2. Solve for the parameters using maximum likelihood estimation

[0223]

[0224] Among them:

[0225] · The parameter values obtained by maximum likelihood estimation, that is, the best parameter estimates obtained by optimizing the likelihood function.

[0226] · β: The parameter to be estimated in the model, usually the regression coefficient or other model parameters.

[0227] · P(Y i |X i , β): Given the input variable Xi Under the conditions of and parameter β, the target variable Y i Conditional probability density function or probability mass function of

[0228] · log P(Y i |X i , β): Take the logarithm of the likelihood function (i.e., the log-likelihood function), which helps simplify calculations and enhance numerical stability.

[0229] · N: The number of samples, representing the total number of samples in the dataset.

[0230] F4.3. Training Process

[0231] 1. Use the Gradient Descent method to optimize the parameters. In the insurance field, the Gradient Descent method can effectively optimize the regression coefficients for predicting insurance claim frequencies or claim amounts, such as the parameters in a Generalized Linear Model (GLM). For example, when predicting the claim amount of a single customer, the model iteratively updates the parameters through the Gradient Descent method to ensure accurate prediction of each customer's claim expenditure.

[0232] 2. Combine Cross Validation to optimize the model hyperparameters, such as the regularization parameter. It divides the dataset into multiple subsets (usually k subsets) and trains and tests the model multiple times to evaluate the generalization ability of the model. During each training process, the selection of hyperparameters can be optimized through cross-validation. For example, through cross-validation, the best regularization coefficient can be selected to prevent the model from overfitting the noise in the insurance data and improve the prediction accuracy on unseen samples.

[0233] 3. Calculate the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC) to select the optimal model. AIC and BIC evaluate the model by balancing the goodness of fit and complexity of the model, avoiding overfitting (the model is too complex) or underfitting (the model is too simple). In claim prediction in the insurance field, by comparing the AIC and BIC values of different models (such as Poisson regression, Gamma regression, etc.), a model that can optimally predict insurance claim situations can be selected. For example, by calculating and comparing the AIC and BIC values of each model, a model that makes the best trade-off between fitting accuracy and model complexity can be selected to ensure that the model can effectively predict and avoid unnecessary computational overhead.

[0234] F4.4. Prediction Optimization After the GLM training is completed, perform prediction optimization:

[0235] 1. Bias Correction: Adjust the model bias based on residual analysis to improve prediction accuracy. In the insurance field, residual analysis can help identify errors in the model for specific customer groups or specific types of claims, thus adjusting the model. For example, in the prediction of insurance claim amounts, there may be cases where certain groups (such as high-risk customer groups) are underestimated. Through bias correction, the prediction results for these groups can be adjusted to make the model more accurately reflect the actual risks of these customers.

[0236] 2. Risk Segmentation: Divide customer risk levels according to the prediction results. By analyzing customers' historical claim data, personal information, and predicted claim probabilities, customers can be carefully segmented by risk, helping insurance companies determine the risk levels of different customer groups. For example, based on the predicted claim amounts of customers, insurance companies can classify customers into low-risk, medium-risk, and high-risk levels. Customers with different risk levels will be assigned different insurance products and pricing strategies to better cope with potential claim risks.

[0237] 3. Dynamic Pricing: Adjust personalized premiums according to risk levels. As customers' risk levels change, insurance companies can adjust premiums in real time to ensure the fairness of premiums and the financial stability of insurance companies. For example, for high-risk customers (such as those with a higher probability of claims), insurance companies can increase their premiums to cover potential payout costs. For low-risk customers, the premiums may be lower, thus motivating customers to purchase insurance and improving customer satisfaction and loyalty.

[0238] Table 1: Example Optimization Strategies

[0239] Risk level Predicted claim rate Rate adjustment Low risk <0.3 Premium decreased by 5% Medium risk 0.3-0.8 Remain unchanged High risk >0.8 Premium increased by 10%

[0240] G. An information disclosure and data security control method G0 designed for the insurance industry is implemented as follows: In order to meet the regulatory requirements of the government for the transparent utilization of insurance enterprise data and the prediction process while ensuring data security, privacy, and compliance, this method designs an information disclosure and data security control mechanism. On the basis of fully encrypting and storing insurance data, an interface is reserved for government regulatory departments so that, on the premise of complying with data privacy protection, necessary transparent information can be provided to the government.

[0241] The specific implementation steps are as follows:

[0242] 1. Data Encryption Storage: To ensure that data is not illegally accessed or leaked during storage, all sensitive insurance data needs to be encrypted and stored. Assume the insurance data set is Each data item d i may contain sensitive content such as the user's personal information, insurance contract terms, claims settlement records, etc. The data encryption process can be expressed as:

[0243]

[0244] where Enc(d i , k) is the encryption function, k is the encryption key, is the encrypted data item. The encrypted data is stored in the insurance company's data center and provides an authorized access method externally.

[0245] 2. Privacy Protection & Anonymization: To meet the requirements of privacy protection regulations, especially when dealing with personal information, this method designs a data anonymization mechanism. Assume the data contains personal identification information (PII). By using data de-identification techniques (such as k-anonymity), sensitive information is replaced or encrypted, thereby removing the identification marks and making the data transformed into a dataset without personal identification information

[0246]

[0247] Data anonymization ensures that personal data is not leaked, so that data analysis and processing can still be carried out while complying with privacy protection.

[0248] 3. Government Interface & Transparent Data Disclosure: To meet the government's need for transparency in the insurance prediction process, this method designs a secure interface that allows government regulatory agencies to access the encrypted stored data through authentication and conduct necessary data audits and monitoring. The implementation of this interface provides a set of controlled access rights and operation rights for the government, enabling the government to access specific model prediction results and data without disclosing the enterprise's sensitive information. Assume the government interface is g, then the access process of the interface can be expressed as:

[0249]

[0250] where, represents the government access interface, Authenticate(gov_key) is the government authentication process, and gov_key is the key or certificate provided by the government, It is a transparent data view that can be accessed by the government. This data view only contains non-sensitive information, ensuring that the government can audit and regulate the data usage of the insurance industry while protecting personal and business confidential information.

[0251] 4. Transparency of the Data Prediction Process: The data prediction process in the insurance industry is crucial for government supervision, especially when it comes to risk assessment and pricing models. This method makes the government able to audit data processing and model inference by making part of the prediction model public and transparent, while protecting the business secrets of insurance companies. Specifically, the intermediate results in the model training and prediction processes (such as model parameters and predicted decision boundaries) can be shown to the government through encryption or "interpretability" techniques (such as SHAP or LIME):

[0252]

[0253] Among them, is the insurance data set and the prediction process under the model parameters θ, is the prediction result that the government can view. In this way, insurance companies can provide the government with sufficient information without disclosing the model parameters, ensuring the fairness and transparency of the prediction process.

[0254] 5. Data Monitoring & Auditing: Insurance companies need to conduct regular internal and external audits to ensure that data processing and storage comply with relevant laws and regulations and there is no unauthorized data access. To achieve this, this method designs a real-time data monitoring and auditing mechanism to record all access requests, data modifications, and prediction operations. All audit logs will be encrypted and stored in a dedicated audit database for subsequent review and analysis. The recording formula of the audit logs is:

[0255]

[0256] Among them, is the audit log recorded at time point t, action represents the operation type (such as data access, prediction request, execution of method A, etc.), timestamp is the time when the operation occurred, user_id is the user identifier who executed the operation, and data_accessed is the data item accessed. These logs provide a traceable audit record for the government and insurance companies to ensure the legality of data usage.

[0257] 6. Disaster Recovery & Data Recovery: To address data loss or tampering caused by emergencies, this method designs a disaster recovery and data recovery mechanism. During the storage of insurance data, all data copies will be regularly backed up, and the backup data will be stored in different data centers with geographically dispersed locations. The data recovery process can be expressed as:

[0258]

[0259] where, is the recovered data, is the original data, and backup_location is the location where the backup data is stored. Through this mechanism, in case of hardware failures or data loss, the insurance company can quickly recover the data and continue to ensure the normal operation of the business.

[0260] Through the above mechanism, insurance enterprises can ensure the privacy of customer data under strict data security control and meet the government's requirements for the transparent utilization of insurance data. While meeting the requirements of privacy protection and compliance, this method ensures that the data processing process in the insurance industry has sufficient transparency and audit capabilities, providing the necessary technical support for government supervision.

[0261] H. An intelligent insurance data processing and decision-making optimization method H0 based on large language models and human expert feedback is specifically implemented as follows:

[0262] 1. Preprocessing of multi-source heterogeneous data: The method of the present invention first performs preprocessing of multi-source heterogeneous data. Specifically, the data includes, but is not limited to, structured data (such as customer information, historical claim records, etc.) and unstructured data (such as customer reviews, social media information, etc.) from the insurance company's database. To process this data, the present invention uses a large language model to clean, denoise, fill in missing values for the text data, and convert it into a structured format to ensure the consistency and integrity of the data.

[0263] The mathematical formula is as follows:

[0264]

[0265] where, represents the original data set, represents the processed data set, f is the data processing function, and LLM represents the large language model.

[0266] For the processed data, as the basic data for the production process of the present invention, it is used in all processes of method F.

[0267] 2. Intelligent Decision Support during Model Training: During the model training process, intelligent decision support is provided with the help of large language models. The present invention optimizes the model training process in the following two aspects:

[0268] Dynamic Hyperparameter Tuning: The large language model automatically optimizes hyperparameters (such as learning rate, batch size, etc.) based on real-time feedback during the training process, accelerating the training and improving the model performance.

[0269] Business Logic Optimization: The large language model combines the current business requirements to optimize the business logic of the model, ensuring that the model output conforms to the insurance business rules.

[0270] In this way, the present invention can effectively improve the intelligence level of the training process and automatically adjust the training strategy according to the feedback results.

[0271] The mathematical formula is as follows:

[0272]

[0273] Wherein, is the feedback during the training process, is the optimization strategy provided by the large language model.

[0274] For the model after optimization iteration, continuously select the best, and use the optimal version of the model as the basis for the preprocessing of the aforementioned multi-source heterogeneous data.

[0275] 3. Prediction Evaluation and Optimization Feedback: In the prediction stage, the present invention uses the large language model and human expert feedback to optimize and feedback the prediction results to ensure the long-term effectiveness of the model. The specific steps are as follows:

[0276] For the feedback opinions provided by the large language model:

[0277] Prediction Result Analysis: The large language model can analyze the results of the model prediction, identify possible errors, and adjust the prediction results.

[0278] Real-time Feedback and Adjustment: Based on the feedback of real-time data and historical data, the large language model continuously optimizes the prediction strategy to improve the prediction accuracy.

[0279] Business Decision Support: Based on the optimized prediction results, the large language model provides strategy suggestions for decision-makers to ensure the scientificity and executability of the decisions.

[0280] The mathematical formula is as follows:

[0281]

[0282] Wherein, is the original prediction result, is the optimized prediction result.

[0283] Take the prediction results as part of the reference, comprehensively analyze them with the feedback information provided by human experts, and evaluate the results of data applications such as optimized prediction.

[0284] I. A low-cost distributed cluster architecture I0 is specifically implemented as follows:

[0285] Hardware components:

[0286] 1. Audit module D0, independent of other modules, independently audits all business data in the storage server in the cluster, and provides an audit report on a periodic basis.

[0287] 2. Compute servers, using the x86 architecture, select common NVIDIA models such as A100, L40S and other models of graphics cards; the processors select common AMD models such as 7763, 9654 and other models of dual-way CPUs, with 1024GB of memory; the chassis select brands such as ASUS, Supermicro, Gigabyte, etc.; at the same time, install RJ45, QSFP+, IB three-mode network cards; install 4 U.2 7.62T nvme SSD hard drives. The hard drive stores the computer program ICODE. When the computer program ICODE is executed by the CPU, the compute server implements the method F0 described in F.

[0288] 3. Network switches, according to the network card selection of the compute servers, install RJ45, QSFP+, IB switches at the same time.

[0289] 4. Storage servers, using the x86 architecture, select MegaRAID LSI series array cards, 32 16T mechanical hard drives to form a Raid5 array, the processors select common AMD models such as 7763, 9654 and other models of dual-way CPUs, with 1024GB of memory; the chassis select brands such as ASUS, Supermicro, Gigabyte, etc.; at the same time, install RJ45, QSFP+, IB three-mode network cards, and the computer program ISTORE. When the computer program ISTORE is executed by the CPU, the compute server implements the method E0 described in E.

[0290] 5. Data exchange servers, using a lower configuration, select low-cost devices, mainly to meet the needs of government supervision. It stores the computer program G1. When the computer program G1 is executed, the processor executes the information disclosure and data security control method designed for the insurance industry described in G to meet the needs.

[0291] 6. Infrastructure such as external three-phase power supply, monitors, input / output devices, external networks, etc.

[0292] Network design:

[0293] 1. RJ45 Switch: IP address range 192.168.1.2 - 192.168.1.255, gateway 192.168.1.1, subnet mask 255.255.255.0, DHCP enabled, access to external network, taking into account the BMC ports of each device.

[0294] 2. QSFP+ Switch: IP address range 192.168.20.2 - 192.168.20.255, no gateway, subnet mask 255.255.255.0, DHCP not enabled, no access to external network.

[0295] 3. IB Switch: IP address range 192.168.30.2 - 192.168.30.255, no gateway, subnet mask 255.255.255.0, DHCP not enabled, no access to external network.

[0296] Storage Architecture Design:

[0297] 1. Compute-in-Memory: The SSD hard drives installed in each computing server jointly form a distributed storage system. Method E0 is selected for this system to improve the computing read / write I / O efficiency while balancing the network load.

[0298] 2. The storage server performs full offline backups of business data regularly. The backup methods include full backups and incremental backups.

[0299] User Interaction Design:

[0300] I. For Method F, the interface design includes the following main modules:

[0301] 1. Data Upload Module: Allows users to select local insurance data files (such as CSV, Excel formats) by clicking the "Upload File" button, and automatically completes data format checking and preprocessing prompts to ensure that the data meets the model input requirements.

[0302] 2. Data Preprocessing Module: Displays steps such as data cleaning, missing value handling, and outlier detection through visualization tools. Users can selectively view and adjust the preprocessing results. For example, users can manually adjust the missing value filling strategy or automatically apply standardization and normalization operations.

[0303] 3. Feature Engineering Module: This module provides interactive feature screening and transformation functions. Users can select feature variables by clicking on the interface or perform feature transformations according to business requirements, such as logarithmic transformation, one-hot encoding, etc. For complex data partitioning operations, the system automatically prompts and recommends the best partitioning strategy (such as a time-based test set). Model Selection and Configuration Module

[0304] This module provides a user-friendly model selection interface, supporting the automatic recommendation of appropriate GLM model variants based on the characteristics of the target variable. Users can choose:

[0305] Poisson regression: used to predict claim frequencies. Gamma regression: used to predict claim amounts. Tweedie regression: handles mixed data of zero claims and positive indemnity data. Users can also customize model configurations, such as selecting appropriate regularization parameters, setting hyperparameter optimization strategies, etc. The interface provides real-time feedback, and users can flexibly adjust parameters through elements such as sliders and input boxes and view real-time simulation results.

[0306] 4. Training and Evaluation Module

[0307] During the training and evaluation phases, users can initiate the training process and view key metrics such as the training status, loss function values, and number of iterations in real time. Specific functions include:

[0308] Training status feedback: The system provides a real-time progress bar and logs, showing the current training round, loss changes, and training time.

[0309] Cross-validation configuration: Users can select a cross-validation strategy (such as k-fold cross-validation) and view performance metrics for each fold (such as accuracy, AUC value, etc.).

[0310] Model evaluation: After training is completed, the system provides various evaluation metrics (such as AIC, BIC, bias-corrected results, etc.). Users can view information such as the model's goodness of fit and complexity to assist in decision-making for selecting the optimal model.

[0311] 5. Prediction and Optimization Module

[0312] After completing model training, users can use this module to predict and optimize insurance data. The system provides the following functions:

[0313] Prediction result display: Users can select the dataset to be predicted (such as new insured customers, historical data, etc.). The system automatically generates prediction results and displays the comparison between the predicted values and the actual values through visualization charts (such as bar charts, scatter plots, etc.).

[0314] Risk assessment and grading: Based on the prediction results, the system automatically generates a risk assessment report and classifies customers by risk level. Users can view the risk assessment data for each customer and provide customized insurance recommendations for different risk levels.

[0315] Optimization suggestions: By analyzing the prediction bias, the system automatically recommends bias correction strategies to help users adjust model parameters to optimize the prediction results. Users can also adjust insurance strategies or business processes according to the optimization suggestions.

[0316] 6. Interaction Feedback and Reporting Module

[0317] This module provides users with detailed interaction feedback and reporting generation functions:

[0318] Report Generation: Users can choose to export the current operation process and results as PDF or Excel reports. The report content covers data preprocessing steps, model selection and configuration, training evaluation results, prediction results, and optimization suggestions, etc.

[0319] Real-time Feedback: The system gives prompts based on real-time data feedback of user operations. For example, when users set unreasonable parameters, the system will guide users to make adjustments through pop-up prompts or warning messages to ensure the accuracy of the operation process.

[0320] Customized Notifications: Users can set notification reminder functions for important operations. For example, when training is completed, prediction results are generated, or reports are available for download, the system will inform users via email or push notifications.

[0321] II. For Method G, the interface design includes the following main modules:

[0322] 1. Login and Authentication Interface

[0323] Before entering the system, users (such as insurance company employees, government regulators, etc.) need to go through the authentication process. The login interface provides input boxes for username and password and supports two-factor authentication (such as SMS verification code or identity authentication device) to ensure that only authorized personnel can access the system. The interface is simple and intuitive, and automatically jumps to the main interface after successful authentication.

[0324] 2. Data Encryption Storage Settings

[0325] Users can configure encryption storage options through the system settings interface. By default, the system encrypts and stores all sensitive data, and users can choose different encryption algorithms and key management strategies. The interface provides detailed descriptions of encryption schemes and operation guidelines to ensure that users can reasonably configure encryption policies while following security specifications.

[0326] 3. Data Privacy Protection and Anonymization

[0327] When processing data, users can choose whether to enable the data anonymization function, and the system automatically detects and prompts which data belongs to personally identifiable information (PII). During the anonymization process, users can view data changes in real time and can set the intensity and method of anonymization (such as k-anonymity or de-identification). The user operation interface provides a detailed description of the privacy protection policy and requires users to confirm that the anonymization process has been completed.

[0328] 4. Government Interface and Data Disclosure Management

[0329] Government regulatory authorities access the data stored by insurance companies through authentication. To ensure transparent disclosure, users need to preset interface access permissions in the system. The system allows setting the time window for government access, the scope of accessed data, and whether to disclose specific model prediction results. A graphical permission management tool is provided in the interface, allowing users to easily set access permissions for different roles to ensure compliant and controlled data disclosure.

[0330] 5. Prediction Transparency View

[0331] To meet government regulatory requirements, users can view and manage the transparency of the prediction process of insurance data through the system. This interface will display some processes and intermediate results of the prediction model. Users can choose whether to disclose key contents such as model parameters and prediction decision boundaries. The interface provides concise interpretability tools (such as SHAP, LIME, etc.) to help users understand the decision-making process of the model and ensure that information disclosure does not leak business secrets.

[0332] 6. Audit and Data Monitoring Logs

[0333] Users can view all operation records in the audit log interface, including data access, modification, and prediction requests, etc. The logs will be automatically classified according to time, operation type, and operator, and support export on demand. The log display method in the interface supports real-time monitoring and allows users to set notification warnings to promptly detect and respond to abnormal operations.

[0334] 7. Disaster Recovery Settings and Data Backup

[0335] Users can set the data backup period, backup storage location, and recovery process through the disaster recovery management interface. The system supports automated backup and allows users to view the backup status in real time. The data recovery process is simple and intuitive. Users only need to select the backup version and confirm the recovery operation, and the system will automatically execute the data recovery process to ensure continuous business operation.

[0336] 8. Data Security Policy and Compliance Check

[0337] In the system settings, users can view and configure data security policies. The system regularly performs compliance checks to ensure that all operations comply with privacy protection regulations and security standards. The check results will generate a detailed report, and users can adjust the policies according to the report content to ensure that data processing complies with legal and regulatory requirements.

[0338] III. For Method H, the interface design includes:

[0339] 1. Data Preprocessing Interface

[0340] In the first step of the system, users can see a clear multi-source data import interface. Users can upload various types of data files, including structured data (such as customer information, claims records, etc.) and unstructured data (such as social media comments, customer feedback, etc.).

[0341] Data import: Users select and upload data through the "Upload" button, and the system supports the upload of formats such as CSV and Excel.

[0342] Data preview and cleaning: After users upload data, the system automatically presents a preview of the data and marks the missing values, outliers, and noisy parts in the data. Users can choose to clean manually or use the "Intelligent Cleaning" function. The system will perform automatic cleaning based on the large language model, fill in the missing values, and denoise the text data.

[0343] Format conversion and output: The cleaned data is automatically converted into a unified structured format. Users can download the processed data file or directly enter the next model training interface.

[0344] 2. Intelligent decision support interface (optimization of the training process)

[0345] The interface in the model training stage is designed to provide real-time feedback and automatic optimization functions.

[0346] Training progress and parameter monitoring: The interface shows the real-time progress of model training, including indicators such as the current round, training loss, and prediction accuracy. Users can view the detailed parameters of the training process, adjust, or optimize the hyperparameters.

[0347] Dynamic hyperparameter adjustment: The large language model will automatically adjust the hyperparameters (such as learning rate, batch size, etc.) according to the real-time feedback, and prompt the user with the current optimized hyperparameter settings through pop-up windows or notifications to ensure more efficient training.

[0348] Business logic optimization suggestions: Combining the current business requirements, the system will propose model adjustment suggestions, and users can choose to accept or adjust manually. The interface contains a set of business rules (such as insurance product classification, compensation standards, etc.). The large language model will automatically optimize the model output according to these rules to ensure that the prediction results conform to the business logic.

[0349] 3. Prediction evaluation and optimization feedback interface

[0350] In the prediction step, the system helps users optimize the prediction results through the intelligent feedback function and supports the integration of expert opinions.

[0351] Prediction result display: Users can view the original prediction results of the model. The system displays the error between the predicted value and the actual value in the form of charts, lists, etc., and points out the possible deviations.

[0352] Intelligent Feedback and Adjustment: Based on real-time data and historical feedback, the large language model automatically analyzes the prediction results and provides optimized prediction results on the interface (e.g., adjusted risk assessment values, customer demand forecasts, etc.). Users can choose to directly adopt the optimized results or click the "Expert Feedback" button to view the adjustment opinions put forward by human experts.

[0353] Decision Support: Based on the optimized prediction results, the system provides a series of strategic suggestions for decision-makers. The system analyzes the matching degree between the optimized data and the business scenario, gives the best decision-making path, and ensures that the prediction results are not only scientific but also executable.

[0354] 4. Integrated Feedback and Expert Advice Interface

[0355] This interface provides a module for integrating human expert feedback. Users can view the specific feedback opinions provided by experts and compare them with the prediction results of the model.

[0356] Expert Feedback Input: Experts can input modification suggestions for the prediction results through the "Feedback" button. The system automatically records the expert's feedback content and updates the corresponding prediction results. Comprehensive Analysis and Optimization: The system comprehensively analyzes the model prediction results and expert opinions, automatically generates an optimized data model, and displays it to users in the form of an "Optimization Report". Users can view the effects of different optimization schemes and select the most suitable one according to actual needs.

[0357] 5. Model Training and Prediction History

[0358] During each training and prediction process, the system generates detailed records, and users can view the historical data and optimization process at any time.

[0359] Historical Record Query: Users can query by conditions such as date, model version, prediction type, etc., view the results of past training or prediction, and ensure that they can trace any decisions and results in the optimization process.

[0360] Result Export and Report Generation: The system supports exporting the results of training and prediction as reports or data files. Users can download a detailed report on the model optimization process for subsequent reference or decision-making basis.

[0361] The present invention provides a distributed clustered model training device for insurance big data, aiming to solve the deficiencies of the prior art in terms of computing efficiency, stability, security, and intelligence. The device combines distributed computing, cloud computing architecture optimization, industry data security control, and large language model-assisted strategies to enhance the overall ability of the insurance industry in big data processing and intelligent modeling. The main technical effects include the following aspects:

[0362] Improve the processing and training efficiency of large-scale heterogeneous data. In view of the complexity of big data in the insurance industry, the present invention designs an efficient data preprocessing and training framework, which supports accurate modeling of multi-dimensional and highly heterogeneous data. Through a distributed computing architecture, the device can efficiently split and process large-scale data in parallel, thereby greatly improving the efficiency of data cleaning, feature extraction, and model training. Compared with traditional single-machine or small-scale cluster training methods, the device has achieved significant optimization in data throughput, training speed, and parallel computing ability, ensuring that model training can meet the needs of massive insurance data.

[0363] Optimize the cloud computing architecture and reduce resource consumption and costs. The insurance industry has limited computing resource budgets. The present invention optimizes the design of the cloud computing architecture and adopts a low-cost computing resource scheduling strategy to maximize the utilization rate of computing resources. Through dynamic task scheduling and distributed resource pool management, the device can be elastically expanded according to actual computing needs, reduce waste of computing resources, and achieve maximum cost-effectiveness. Experimental results show that, compared with traditional cloud computing architectures, the device can improve computing efficiency by about 30%-50% under the same computing resource conditions and effectively reduce operating costs.

[0364] Strengthen data security and industry compliance. The insurance industry has extremely high requirements for data security and compliance. The present invention adopts a multi-level security mechanism in the process of data storage and transmission, including means such as data encryption, access control, and data desensitization, to ensure privacy protection and compliance of insurance data during the training process. At the same time, the device is built-in with a security module that meets industry regulatory requirements and can meet the relevant regulations of the insurance industries of various countries (such as GDPR, CCPA, etc.), ensuring the safe use and compliant storage of data.

[0365] Introduce large language models to improve the degree of intelligence. Traditional insurance data modeling relies on manual feature engineering or static rule setting and is difficult to dynamically adapt to business needs. The present invention introduces large language models into the training device to provide intelligent auxiliary support, covering multiple links such as data cleaning, feature extraction, and model optimization. Experiments show that during the insurance data modeling process, the device can reduce the workload of manual feature engineering by 40% and improve the accuracy of the prediction model, effectively enhancing the intelligence level of the business model.

[0366] Enhancing Human-Machine Collaboration and Improving Model Adaptability By innovatively introducing a human expert feedback mechanism, the present invention enables experts to intervene and adjust in real time during the model training process, improving the training effect. The combination of experts' industry experience and data-driven model training makes the final model not only have stronger prediction ability but also better meet the actual needs of the insurance business scenario. Compared with the completely automated training method, the adaptability of the present invention in insurance industry data modeling has been improved by more than 20%, significantly enhancing the usability and business adaptability of the model.

[0367] In summary, through distributed computing, large-scale data processing, cloud computing optimization, security mechanism strengthening, intelligent strategy introduction, and human-machine collaboration optimization, the present invention has significantly improved the intelligent model training ability of the insurance industry in the big data environment, meeting the core requirements of the insurance industry for efficiency, security, and intelligence. Brief Description of the Drawings

[0368] Figure 1 It is a relationship diagram between various parts of the present invention. Detailed Description of the Invention

[0369] The following combines the drawings and specific embodiments to elaborate on the present invention. In this specification, the drawing size ratio does not represent the actual size ratio. It is only used to reflect the relative positional relationship and connection relationship between components. Components with the same name or the same reference numeral represent similar or identical structures, and are for illustrative purposes only.

[0370] For the data causality check module A0 covering the entire production chain, the embodiment is as follows: This embodiment uses multi-dimensional data including the basic information of the insured, vehicle information, historical claim records, repair behaviors, and social network relationships. The data samples are shown in Table 2.

[0371] Table 2: Insurance-related data samples of the insured

[0372]

[0373] Causality Check Analysis

[0374] Variable Correlation Analysis

[0375] First, use the Pearson correlation coefficient to analyze the relationship between variables and the claim rate:

[0376]

[0377] The analysis results are as follows:

[0378] · Driving experience and claim rate: ρ = -0.42 (negative correlation, the longer the driving experience, the lower the claim rate)

[0379] Credit score and claim rate: ρ = -0.38 (the higher the credit score, the lower the claim rate)

[0380] Vehicle age and claim rate: ρ = 0.25 (vehicles with older vehicles have a slightly higher claim rate)

[0381] Maintenance behavior (frequent use of the same repair shop) and claim rate: ρ = 0.48 (strong positive correlation)

[0382] Causal Inference Analysis

[0383] Create a causal graph:

[0384] Driving habits → accident rate → compensation frequency

[0385] Maintenance habits → accident rate → compensation amount

[0386] Use Do-Calculus to estimate causal effects:

[0387]

[0388] Among them, if the maintenance plant selection variable is controlled:

[0389] P(claim rate|repair action=no)=0.15 (10)

[0390] P(claim rate|repair behavior=yes)=0.42 (11)

[0391] Counterfactual reasoning: If policyholders with poor driving habits (such as frequent speeding and sudden braking) begin to comply with safe driving regulations, the claim rate will drop by 10%; if policyholders whose vehicles have been in use for more than 5 years perform regular maintenance, the claim rate will drop by 12%; if policyholders with a long history of traffic violations reduce the number of violations, the frequency of compensation will drop by 15%.

[0392] Calculation of variable influence weights: The SHAP method was used to evaluate the variable influence weights, with driving habits accounting for 40%, brand accounting for 20%, vehicle characteristics accounting for 15%, maintenance behavior accounting for 10%, agent factors accounting for 8%, and social network relationships accounting for 7%.

[0393] Conclusion and Optimization Strategies

[0394] **Customer characteristics with high claim rates**: driving experience < 10 years, credit score < 700, history of traffic violations, frequent use of specific repair shops.

[0395] **Optimization measures**:

[0396] - Increase premiums by 5%-10% for high-risk customers;

[0397] – Introduce a driving behavior improvement plan, such as safety driving rewards;

[0398] – Monitor the behavior of repair shops to reduce moral hazards.

[0399] Dynamic Causal Diagram Network Modeling

[0400] In the field of auto insurance, the present invention considers a dynamic causal diagram model to capture the causal relationships that change over time among policyholders, vehicles, and their historical behaviors. This model not only considers static causal relationships but also the dynamic evolution of causal relationships over time. The specific modeling process is as follows:

[0401] Modeling idea: The dynamic causal diagram aims to describe the causal structure that evolves over time and is suitable for modeling time series data. Different from the static causal diagram, the nodes and edges in the dynamic causal diagram are updated with time steps, enabling it to better capture the impact of changes in different factors over time on the result variable.

[0402] In this model, the following variables are mainly considered:

[0403] · Driving habit (D t ): Describes the driving behavior of the policyholder at time point t, such as whether they often speed or brake suddenly;

[0404] · Violation record (V t ): Describes the violation situation of the policyholder at time point t;

[0405] · Repair behavior (M t ): Describes whether the policyholder frequently uses the same repair shop or conducts regular vehicle maintenance;

[0406] · Claim frequency (C t ): The claim frequency of the policyholder at time point t;

[0407] · Claim amount (P t ): The claim amount of the policyholder at time point t;

[0408] The construction of the dynamic causal diagram is based on the following core assumptions:

[0409] · The driving habit of the policyholder affects the violation record, and the violation record further affects the claim frequency;

[0410] · The repair behavior directly affects the claim frequency and the claim amount;

[0411] · The claim amount is affected by historical behaviors (such as driving habit, repair behavior) and the occurrence of the current accident.

[0412] Dynamic Causal Diagram Structure

[0413] In the dynamic causality diagram, the nodes at time step t represent the behavior of the policyholder and the claim situation. At each time step, the state of the node is affected not only by the previous time step (t - 1), but also by its own behavior at the current time step.

[0414] The structure of the dynamic causality diagram is as follows:

[0415] D t →V t →C t →P t

[0416] M t →C t →P t

[0417] In this model:

[0418] ·D t (Driving habit) directly affects V t (Violation record), and the violation record has a direct impact on C t (Claim frequency);

[0419] ·M t (Maintenance behavior) has a direct impact on C t (Claim frequency), and at the same time, the maintenance behavior also affects P t through C t (Claim amount);

[0420] ·There is a direct causal relationship between the claim frequency C t and the claim amount P t .

[0421] In addition, the dynamic causality diagram also allows the present invention to adjust the model according to historical data and infer the causal effects in different scenarios. For example, by intervening in the driving habit (such as implementing a safe driving reward), the present invention can predict its impact on future claim frequency and claim amount.

[0422] Dynamic causal inference

[0423] To perform dynamic causal inference, the present invention needs to perform a temporal modeling of the above causality diagram. The present invention assumes that at each time point, the state of the system is described by the following conditional probability distribution:

[0424] P(D t ,V t ,C t ,P t |D t-1 ,V t-1 ,C t-1 ,P t-1 )

[0425] Among them, D t represents the driving habit at the current time step, V t represents the violation record at the current time step, C t represents the claim frequency at the current time step, P t represents the claim payment amount at the current time step. The conditional probability distribution can be learned from historical data.

[0426] Counterfactual Reasoning and Intervention

[0427] Based on the dynamic causal graph model, the present invention can perform counterfactual reasoning to evaluate the effect of intervention. For example, the present invention can simulate the following scenarios:

[0428] · If the policyholder improves their driving habit (e.g., reduces violation behaviors), how will the future claim frequency and claim payment amount change?

[0429] · If the policyholder starts regular vehicle maintenance, can it effectively reduce the claim payment amount?

[0430] By intervening in the causal graph (e.g., changing D t to "safe driving"), the present invention can predict and quantify the effect of the intervention measure. This method is of great significance for formulating insurance policies, evaluating risks, and optimizing claim settlement strategies.

[0431] Embodiment: Causal Logic Reasoning Process and Dynamic Deduction Process of Insurance Data

[0432] This embodiment combines data analysis in the insurance industry to show how to optimize the model and make reasoning decisions based on causal reasoning and dynamic deduction processes, especially their applications in data auditing and risk assessment. The process is based on the theory of contradiction and Socrates' syllogism, and designs a logical reasoning and deduction method for insurance data.

[0433] Causal Logic Reasoning Process

[0434] Background: In the insurance industry, causal reasoning is widely applied in customer behavior analysis, claim frequency prediction, potential risk identification, etc. Through causal reasoning, it can help insurance companies understand the causal relationships between different variables, so as to make more accurate decisions.

[0435] Steps of Causal Reasoning

[0436] The causal reasoning process can be achieved through the following steps:

[0437] 1. Define variables: Determine the causal relationship variables in the insurance data, such as "customer age", "vehicle type", "accident history", "claim payment amount", etc. Set the target variable as "claim payment amount", that is, analyze the influence of other factors on the claim payment amount.

[0438] 2. Construct a causal relationship network: Construct a graph data structure, representing different variables as nodes and the causal relationships between variables as edges. For example, there may be a causal relationship between the nodes "customer age" and "historical accident records", indicating that older customers may have more accident histories. Similarly, there may be a direct causal relationship between "historical accident records" and "claim amount".

[0439] 3. Identify causal chains: After constructing the graph model, identify potential causal chains. For example, assume that "customer age" affects "accident history", and "accident history" in turn affects "claim amount". Therefore, the entire causal chain can be represented as:

[0440] Customer age → Accident history → Claim amount

[0441] 4. Causal effect estimation: Based on the constructed causal network, use causal inference methods (such as counterfactual reasoning) to estimate the effect of a certain causal relationship. For example, assume that we want to know how increasing the claim history of customers in a specific age group affects the claim amount. In this case, counterfactual reasoning can give the corresponding claim amount estimate.

[0442] 5. Verify the results of causal reasoning: Verify the results of causal reasoning based on actual insurance data. For example, if the model predicts that the claim amount of a certain group will increase, the actual data should verify whether this prediction holds. If the prediction is consistent with the actual situation, it indicates that the reasoning model is reliable; if not, the causal reasoning network needs to be adjusted and the relationships between variables need to be re-evaluated.

[0443] Dynamic deductive reasoning process

[0444] Background: Dynamic deductive reasoning is a process of predicting future events based on existing data and knowledge through time series analysis or reasoning models. In the insurance industry, dynamic deductive reasoning can help companies analyze the development trends of different events and make corresponding decisions.

[0445] Dynamic deduction steps

[0446] The dynamic deductive reasoning process mainly includes the following steps:

[0447] 1. Data input and initial state setting: Insurance companies first collect insurance data, such as "customer ID", "insured amount", "insurance history", "claim amount", etc., and set the initial state. For example, assume that a customer's insured amount is 1000 and the historical claim record is 200.

[0448] 2. Definition of deductive rules: Define deductive rules based on historical data and business rules. For example: "If the insured amount of a customer exceeds 800, the probability of their claim amount in the next 5 years is 0.4". Such rules are derived from actual business data and statistical analysis and are used to predict the future risks of customers.

[0449] 3. Counterfactual reasoning and hypothesis verification: Based on counterfactual reasoning, different hypothetical scenarios can be generated during the dynamic deduction process. For example, assuming that a customer's insured amount increases from 1000 to 2000, counterfactual reasoning will help predict the change in the customer's future claim amount.

[0450] 4. Model deduction: Based on the defined deductive rules and the results of counterfactual reasoning, use an inference model to dynamically predict customer behavior or claim history. This process can simulate the impact of different factors on future claim situations, thereby providing predictive data for insurance companies.

[0451] 5. Application of deductive results: Finally, the results of dynamic deductive reasoning will be applied to decision-making. For example, if a customer has a high insured amount and a large predicted claim amount, the insurance company can decide whether to adjust their premium or risk strategy based on this result.

[0452] Combining the reasoning methods of the Theory of Contradiction and Socrates' syllogism

[0453] Application of the Theory of Contradiction

[0454] The Theory of Contradiction emphasizes promoting the development and change of things by revealing the internal contradictory relationships of things. In insurance data analysis, potential conflicts and risk points can be identified through the Theory of Contradiction. For example, if a customer makes multiple claims in a short period of time, it may conflict with the insurance company's goal of reducing the claim frequency. In such cases, more reasonable decisions can be promoted through the revelation and reconciliation of contradictions.

[0455] Suppose there is the following data: The claim history of customer A shows that they have 4 claim records in the past two years, but each claim amount is relatively small. The insurance company hopes to reduce the claim frequency in the short term to control risks. However, there is a contradiction between customer A's continuous claim history and the relatively low claim amount: Although the claim amount is small, the excessive number of claims may lead to the accumulation of risks. Through the reasoning of this contradiction, it can be concluded that in order to reduce risks, the company needs to introduce a more accurate claim review mechanism.

[0456] Application of Socrates' syllogism

[0457] Socrates' syllogism is a method of drawing conclusions through strict logical reasoning. Its structure is: major premise → minor premise → conclusion. In insurance data analysis, risk points can be identified through syllogistic reasoning.

[0458] · Major premise: All customers with high insured amounts have a high claim frequency.

[0459] · Minor premise: Customer A has a high insured amount.

[0460] · Conclusion: Therefore, Customer A has a high claim frequency.

[0461] This reasoning process can help insurance companies identify potential high-risk customers, thereby adjusting premium or claim settlement strategies.

[0462] Summary

[0463] This embodiment demonstrates how to conduct causal reasoning and dynamic deduction in insurance data analysis by combining the theory of contradiction and the Socratic syllogism. By establishing a causal relationship network, performing counterfactual reasoning, revealing contradictions in the data, and combining the logical reasoning of the syllogism, insurance companies can more effectively predict claim risks, optimize insurance product designs, and make more accurate decisions in a changing market environment.

[0464] Model Evaluation

[0465] To verify the effectiveness of the model, the present invention uses historical insurance data to evaluate it. The present invention evaluates the performance of the model by calculating the bias of the causal effect, the confidence interval, and the accuracy of the prediction results. The evaluation metrics include:

[0466] · Mean Square Error (MSE) between the predicted claim amount and the actual claim amount;

[0467] · Intervention effect of counterfactual reasoning;

[0468] · Robustness and stability of the model.

[0469] Through experiments, the present invention finds that the dynamic causal graph can effectively capture causal relationships and predict risk changes brought about by different factor changes.

[0470] Conclusion

[0471] This section demonstrates how to combine causal inference with time series data through dynamic causal graph network modeling and apply it to the field of auto insurance. By considering the causal relationships of policyholder behavior over time, the present invention can better understand the dynamic changes of insurance risks and provide effective decision support for insurance pricing and risk prediction.

[0472] For the optimization method of dynamic deduction data audit combining virtual and real, the embodiments are as follows:

[0473] The basic process of this method includes the following steps:

[0474] · Collect actual data: including data such as the water leakage situation of the house roof, the construction years of the houses in the block, the historical claim situations in the same block, and the house maintenance records.

[0475] · Build a virtual simulation model: Build a virtual simulation model for the house water leakage problem based on historical data and expert knowledge to simulate the occurrence of water leakage under different conditions.

[0476] · Conduct dynamic deductive reasoning: Combine the virtual model with the actual data to conduct temporal reasoning on house water leakage and identify possible risk points.

[0477] · Counterfactual reasoning and auditing: Introduce counterfactual reasoning into the model to simulate the changes in the water leakage incidence rate under different scenarios (e.g., house maintenance, house reconstruction, etc.), so as to audit the rationality of the data.

[0478] The inputs of the system include the following data columns: house ID (uniquely identifying each house), roof water leakage situation (recording whether the house roof leaks and the leakage frequency), construction years of the houses in the block (recording the construction years of the houses in the block), historical claim situations (recording the historical water leakage or other claim situations of the house), maintenance history (recording the maintenance records of the house, including roof repairs, regular inspections, etc.), climate conditions (climate data in a specific block, such as precipitation, wind speed, etc.), and audit results (recording the findings and review results during the audit process).

[0479] Table 3: Examples of house water leakage data

[0480]

[0481] Virtual simulation modeling and dynamic reasoning

[0482] Based on the actual data collected, this method builds a virtual simulation model to simulate the occurrence of house water leakage. Suppose we take the houses in a certain block as an example and conduct the following dynamic reasoning:

[0483] Construction years → Water leakage risk → Maintenance record → Probability of water leakage occurrence

[0484] Suppose in a block where the construction years of the houses are relatively long (more than 30 years), the water leakage risk increases with the increase of the construction years. At the same time, the historical maintenance records of the houses also have an important impact on the probability of water leakage occurrence. Through the virtual simulation model, we can dynamically reason about the occurrence and development of water leakage in the time dimension.

[0485] For example, suppose a certain house has a construction year of more than 30 years and has not undergone regular maintenance. The model infers that the probability of water leakage in this house is relatively high. At this time, we can simulate the probability of water leakage occurrence in the house under different scenarios:

[0486] When the house is not maintained, the probability of leakage is 0.25; when the house has undergone one roof repair, the leakage probability drops to 0.15; if the house undergoes a full renovation, the leakage probability is further reduced to 0.05.

[0487] Counterfactual Reasoning and Data Auditing

[0488] Based on dynamic reasoning, counterfactual reasoning is used to simulate the outcomes under different interventions or changes. For example, we can evaluate how the incidence of leakage changes if a certain house undergoes different levels of repair.

[0489] Counterfactual Reasoning Scenario 1: If the house had been regularly maintained within five years after construction instead of waiting until 20 years later for repair, how would the leakage risk change? The simulation reasoning shows that the leakage probability would drop from 0.25 to 0.10 after the early repair.

[0490] Counterfactual Reasoning Scenario 2: If the house is located in an area with relatively harsh climate conditions and an additional roof waterproof layer is added, how would the leakage probability change? The simulation results show that the additional waterproof layer can reduce the leakage probability to 0.08.

[0491] Based on the results of these counterfactual reasonings, auditors can evaluate the reasonableness of the existing data and perform data correction or early warning. For example, if the leakage risks of some houses are underestimated in the historical data, counterfactual reasoning will help reveal these potential auditing defects.

[0492] Audit Optimization and Decision Support

[0493] After completing the dynamic reasoning and counterfactual reasoning, this method can provide more accurate decision support for data auditing. By integrating historical data, virtual simulation models, and reasoning results, auditors can more effectively identify inconsistencies or potential problems in the data and provide a basis for future maintenance decisions. For example, if it is found that the leakage risks of some houses are not identified in a timely manner, auditors can put forward repair suggestions based on the reasoning results of the model or recommend a more detailed inspection of the houses.

[0494] Conclusion

[0495] The present invention provides the above-mentioned dynamic deductive data audit optimization method combining virtual and real. By combining actual data with virtual simulation modeling and using dynamic reasoning and counterfactual reasoning, it can perform effective risk assessment and data auditing in complex scenarios such as house leakage. Through counterfactual reasoning, auditors can predict the effects of different intervention measures, optimize the audit process, and improve the accuracy and efficiency of data auditing.

[0496] For the above-mentioned insurance data distributed storage method E0 after graph modeling, the embodiments are as follows:

[0497] 1. Data Sharding (E1): Insurance data is first sharded by the Storage Management Unit (SMU). Specifically, data such as historical claim records, customer information, and policy information is divided according to certain rules (such as region, time, risk type, etc.) and stored on different storage nodes respectively. This can avoid a single storage node from bearing too much data and prevent bottlenecks in the data reading and updating processes. Each storage node is only responsible for storing a part of the data, thus improving the storage efficiency.

[0498] 2. Data Redundancy and Backup (E2): To ensure the security and reliability of data, this method adopts a data redundancy mechanism. Each data shard is backed up on multiple storage nodes to ensure the integrity of the data in case some nodes fail. For example, the data shard 1 on storage node A can be simultaneously copied to nodes B and C to form redundant backups. This redundant backup mechanism can effectively prevent data loss and ensure the fault tolerance of the system.

[0499] 3. Graph Data Structure Storage (E3): After the insurance data is sharded and redundantly backed up, the graph data structure (including node and edge information) is mapped to the distributed storage system. The nodes (such as customers, policies, claim events, etc.) and edges (such as the relationship between customers and policies, the historical association between policies and claims) in the graph network are stored on different storage nodes respectively, and unified access control is performed through the Storage Management Unit.

[0500] 4. Distributed Query and Access (E4): When accessing specific insurance data, the Storage Management Unit (SMU) quickly locates the storage node where the data is located through the distributed query mechanism and performs efficient retrieval. The query request is first routed to the Storage Management Unit, and then according to the data sharding rules, the SMU decides which storage node should process the query request. Since parallel access is supported between storage nodes, the query operation can return results in a relatively short time.

[0501] 5. Data Update and Synchronization (E5): Insurance data may change continuously during storage, such as the update of customer information, insurance records, or claim data. Whenever the data changes, the Storage Management Unit automatically updates the corresponding data shards and ensures data consistency and synchronization. The update operation is carried out through an efficient Data Exchange Bus (DSEB) to ensure the timely synchronization of data between storage nodes. This process supports real-time data updates to ensure that insurance companies can obtain the latest customer information and claim records.

[0502] 6. Data Storage Optimization (E6): To further improve storage performance, the system dynamically optimizes based on the access frequency of data. For example, frequently accessed insurance data (such as claim records of high-risk customers) will be automatically migrated to nodes with faster storage speeds (such as SSD storage nodes), while less frequently accessed data can be stored on lower-cost storage media (such as HDD). Through this data storage optimization mechanism, the system can balance storage costs and performance requirements.

[0503] The present invention aims to improve the risk assessment, pricing decision-making, fraud detection, and audit efficiency in the insurance industry. By improving the traditional GLM, combining GNN to identify the correlation between data, and introducing causal analysis for virtual data simulation and dynamic deductive auditing, the present invention realizes precise and intelligent insurance data processing. To clearly describe the application of the present invention, specific implementation examples are provided below.

[0504] Example: Optimization of Auto Insurance Pricing and Claim Review

[0505] I. Background

[0506] 1. Insufficient Pricing Accuracy: Currently, GLM is used for customer pricing, but the social correlation between customers is not considered, resulting in underestimated rates for some high-risk customers and affecting risk balance.

[0507] 2. Frequent Fraudulent Claims: The traditional claim review process relies on a rule engine and is difficult to detect gang fraud, resulting in some unreasonable claims passing the review and increasing the insurance company's payout costs.

[0508] 3. High Cost of Manual Review: The current review process lacks intelligent means, and a large number of low-risk claims still require manual review, resulting in low processing efficiency.

[0509] II. Data Input

[0510] The data used in this example includes:

[0511] Policyholder information includes Policyholder ID, Gender, Age, Driving Experience, Credit Score, and Traffic Violation History.

[0512] Vehicle information includes Vehicle Brand, Vehicle Model, Vehicle Age, Engine Displacement, and Mileage.

[0513] Historical claim data includes the claim count in the last 5 years and the repair shop.

[0514] External associated data includes the agent ID, the frequent use of the same repair shop, and social network relationships.

[0515] The weight variable includes the earned exposure of auto insurance, which represents the effective underwriting time of the auto insurance contract during the calculation period, measured in policy - years, and is used to standardize the risk measures of different policyholders, with a range between 0 and 1.

[0516] The target variables include the claim frequency, which measures the number of claims made by a policyholder within a certain period of time and reflects their accident risk level, usually calculated based on historical claim records; and the claim severity, which is the average amount of each claim made by the policyholder.

[0517]

[0518] Table 4: Basic Information of Policyholders

[0519] III. Implementation Steps

[0520] Step 1: Calculate individual risks based on GLM

[0521] 1. Calculate the claim frequency, and use Poisson Regression to predict the probability of a policyholder having an accident in the next year.

[0522] 2. Calculate the claim severity, and use Gamma Regression to estimate the expected claim cost.

[0523] 3. Combine the claim frequency and the claim severity to calculate the preliminary premium of the customer.

[0524]

[0525] Table 5: Policyholder Claim and Social Information

[0526] Step 2: Identify high - risk customers using GNN - IAO

[0527] 1. Train GNN - IAO to calculate the social risk score of each customer.

[0528] 2. Identify customer groups of shared contacts, repair shops, and insurance agents, and adjust the premiums for high-risk customers.

[0529] Step 3: Apply VR-DDA for fraud detection and claim review

[0530] 1. Generate a large amount of virtual fraud data, including cases of gang fraud, duplicate claims, false accident reports, etc.

[0531] 2. Optimize the claim review strategy through reinforcement learning (RL) to improve the model's ability to identify fraud claims.

[0532] 3. Mark high-risk claims for in-depth review and automatically approve low-risk claims.

[0533] IV. Implementation Effects The backtest results of the present invention on historical data show that:

[0534] Improve pricing accuracy: Identify social connections, increase the rates of high-risk customers by 12%, and decrease the rates of low-risk customers by 8%. After the rate adjustment, the overall risk level of the insurance company decreases, and the loss ratio drops by 2.3%.

[0535] Optimize fraud detection: Generate fraud data and optimize the review strategy, with the fraud identification rate increasing by 19%. By detecting gang fraud behavior, intercept 17% more suspicious cases than before.

[0536] Improve review efficiency: Conduct intelligent claim review, with the automatic approval rate of low-risk claims increasing by 30%. The manual review of high-risk claims decreases by 26%, and the claim settlement processing speed increases by 30%.

[0537] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0538] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementation in the processFigure 1 means for the functions specified in one or more processes and / or blocks Figure 1 or means for the functions specified in one or more blocks.

[0539] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions in the process Figure 1 means for the functions specified in one or more processes and / or blocks Figure 1 or means for the functions specified in one or more blocks.

[0540] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 means for the functions specified in one or more processes and / or blocks Figure 1 or means for the functions specified in one or more blocks.

[0541] The foregoing is only a description of the preferred embodiments of the present invention and is not intended to limit the scope of the present invention. Without departing from the spirit of the design of the present invention, various modifications and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A distributed clustering model training device based on insurance big data and its application, characterized in that: include: A. Data causality check module A0 covering the entire production chain: audits all data change processes in the entire business process to enhance the relationship between variables from correlation to causality in the entire process of data calculation, prediction and analysis. Specifically, it includes: a graph network enhanced insurance data audit optimization method and a virtual-real dynamic deductive data audit optimization method; B. A hardware acceleration method for graph network-enhanced insurance data audit optimization method B0: This method optimizes the audit process of insurance data by customizing hardware accelerators and graph network enhancement technology. It uses dedicated processing units, GPUs, and graph storage units to improve the efficiency of graph data processing, storage, and optimization calculations, and achieves efficient graph construction, deep learning training, and real-time auditing, ensuring the accuracy of insurance data and the real-time nature of audit results. C. A hardware acceleration method C0 specifically for a dynamic deductive data audit optimization method combining virtuality and reality: The method C0 is designed specifically for dynamic deductive data audit optimization combining virtuality and reality, and combines the needs of virtual simulation and real-time data audit. Through a dedicated processing unit SPU, a data mapping accelerator DMA, a dynamic data fusion module DFM, a multi-dimensional computing unit MPCU and a hardware coordination interface HCI, it can efficiently process complex data in the dynamic deductive process. The method supports real-time data analysis and prediction, and improves audit efficiency and accuracy through hardware acceleration, ensuring efficient fusion and optimization of virtual and real data in the audit process. D. The audit module D0 built on the basis of the method B0 and the method C0 described in B and C, through the hardware components such as the processing unit, the memory, the simulation module, etc., combined with program coordination, realizes efficient data audit optimization and result output of the entire business chain; E. A distributed storage method for insurance data after graph modeling E0, which implements distributed storage and computing integration of insurance data by sharding the insurance data after graph modeling and tightly coupling the storage process with the computing process; F. An insurance data processing method F0 based on a generalized linear model (GLM), combined with the business characteristics of the insurance industry, realizes accurate prediction and optimization processing of insurance data, covering multiple links such as data preprocessing, feature engineering, model training, prediction optimization, etc., and is applicable to, including: property insurance pricing, health insurance risk assessment, life insurance actuarial science, fraud detection; G. An information disclosure and data security control method G0 designed for the insurance industry, which meets the government's regulatory needs for the transparent use and prediction process of insurance companies' data while ensuring data security, privacy and compliance. On the basis of fully encrypting and storing insurance data, an interface is reserved for government regulatory departments so that necessary transparent information can be provided to the government while complying with data privacy protection. H. An intelligent insurance data processing and decision optimization method based on large language model and human expert feedback H0, through large language model and expert feedback, optimizes the preprocessing of multi-source heterogeneous data, decision support in the model training process, and evaluation and optimization of prediction results, so as to improve the efficiency of insurance data processing and decision-making; I. A low-cost distributed cluster architecture I0.

2. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The data causality check module A0 covering the entire production chain includes: auditing all data change processes in the entire business process, and enhancing the relationship between variables in the entire process of calculation, prediction and analysis from correlation to causality, including: A1. Considering that the variables in the insurance industry data calculation process are relatively fixed, and there is a basically stable functional relationship between the variables in terms of their implicit physical, practical and economic meanings, but the functional relationship is not explicit and is often not a polynomial form, we first use a graph network to model the relationship between variables to achieve an enhanced representation of the relationship between each variable, and then use the constructed graph network to conduct a causal assessment audit of the data of the entire production process, including: A11. Use historically verified reliable real data as the initial data for network modeling. The data features include: policyholder ID, gender, age, driving experience, credit score, historical violation record, vehicle brand, vehicle model, vehicle age, displacement, mileage, number of claims in the past five years, repair location, agent ID, whether the same repair shop is frequently used, social network relationships, earned premiums for auto insurance, frequency of claims, and amount of each claim; A12. In the process of insurance data processing, since the variable relationships involved are highly nonlinear, non-explicit, and have implicit economic, physical, and social rules, the insurance data audit optimization method enhanced by a graph network is combined with a customized causal inference algorithm to perform deep abstraction and reasoning on the causal relationship in insurance big data. The insurance data audit optimization method enhanced by a graph network includes: A121. Dynamic graph networks and causal relationship modeling, including: First, each variable in the insurance data is regarded as a node in the graph and embedded through a set of functions f θ (x) embedding the features of the nodes, where x represents a data point in the insurance business. Different from the traditional graph convolution operation, the graph structure update method of the present invention does not rely on a fixed adjacency matrix or a static graph structure, but dynamically adjusts the edge weights between nodes based on an implicit causal model; Set each node v i The characteristic is represented as h i , its update process is described by the following formula: in, Represents node v i The set of adjacent nodes at time t, is the dynamic adjustment coefficient of the connection strength between nodes, which is calculated based on the temporal relationship implicit in the data, historical data and the results of the causal inference model; ψ t (h i (t) ,h i (t-1) ) is a temporal dependency generated based on the causal inference algorithm, capturing the changing trend of node features. t is a trainable parameter, f θ Represents a variety of alternative function types, including polynomial functions, exponential functions, logarithmic functions, and fractions; A122. In order to effectively improve the causal reasoning between nodes in the graph network, the present invention introduces an adaptive causal inference function in the network update process, which can dynamically learn the causal relationship between variables according to the evolution of the data flow; In each update round, the causal reasoning process of the node is performed according to the following recursive formula: Among them, Y represents the target variable (such as the final status of insurance claims), X i is the input feature of the i-th node (variable), C i are conditional variables introduced in the process of temporal and causal reasoning, which describe the influencing factors at the current moment. It is a conditional probability distribution, which indicates the probability of occurrence of the target variable Y given the historical data X and the causal condition C; A123, causal path and conditional independence constraints, uses causal path constraints to build complex dependency networks. For each pair of nodes v i and v j , by calculating their causal path probabilities and performing conditional independence constraints, the connection weights of the edges are further optimized, and the causal paths between nodes are calculated using the following formula: in, represents the set of all possible paths, is a conditional independence constraint, used to control the rationality of the causal path, and k represents the path between nodes; A124. Time nesting and collaborative optimization mechanism, including: Considering the significant impact of time series changes in insurance data on causal reasoning, a collaborative optimization mechanism based on time nesting is used. This mechanism introduces a recursive time gating function G t Perform weighted control on the time dimension to dynamically adjust the impact of time series data on causal reasoning: Among them, G t is a time gating function used to control the degree of dependence of node updates on the causal reasoning path within each time step. t is a trainable parameter; A125. State transition functions and causal reasoning between nodes, including: In the graph network model of the present invention, the state transition between nodes is designed as a highly nonlinear and dynamically changing process, aiming to perform adaptive learning and optimization based on the complexity and implicit causal relationship of insurance big data. The state transition of nodes does not only rely on the simple linear relationship between nodes, but reflects more complex variable change laws by integrating historical data, causal inference and time series information. First, consider each node v in the insurance data i Status i , its change with time step t can be expressed as a recursive process. The change of node state is affected by the intrinsic characteristics of the node and the causal relationship between nodes. A new state transition function is introduced Used to describe the slave node v i To node v j The causal state transition at time t, Use the Markov chain function as the basis function for enhanced training: Among them, h i (t) and h j (t) They are node v i and v j The feature representation at time t is is the node containing the node v i and v j The conditional information of the causal relationship between the two nodes. The state transition function describes the evolution of the causal relationship between the two nodes and can be updated according to the historical information and causal paths at different time steps; A125.

1. State Transitions in Implicit Causal Models In order to more accurately capture the state transition in the causal reasoning process, this paper introduces an implicit causal model to further optimize the dynamic update of node states. Assume that node v i The state transformation is not only affected by the input feature h of the node itself i (t) The influence of the historical state h i (t-1) and external environment C i The state transition function can be expressed as: h i (t+1) =h i (t) +f θ h i (t-1) ,C i Among them, f θ Is a nonlinear mapping function used to represent the node v i The change relationship between the state of time t and time t-1. This function can adapt to complex time series data and potential nonlinear causal relationships by embedding feature information of multiple historical moments; A125.2 Structured Causal Path Constraints and Optimization In order to further improve the reasoning accuracy, a structured causal path constraint is introduced for the causal state transition between nodes. This constraint not only considers the direct causal relationship between nodes, but also includes the high-order dependency path of nodes. For this purpose, a causal path function is defined Used to describe the slave node v i To node v j The causal path of: Among them, γ ijk is the coefficient in the causal path, which represents the node v i To node v j The higher-order influence relationship here refers to the influence relationship based on higher-order reasoning logic relative to the first-order reasoning logic; A125.3 Timing Gating Mechanism and Causal Path Adjustment Considering the important influence of timing factors on causal reasoning in insurance data, the timing gating mechanism can accurately control the timing characteristics of information flow by dynamically adjusting the weights of causal paths between moments. The timing gating mechanism is implemented through the following functions: Among them, G t is a gating function used to adjust the strength of information flow on the causal path at time step t. t The model can adaptively respond to dynamic changes in the time series and map this response to the results of causal reasoning. A125.4 Causal Inference and Final Prediction Through the above-mentioned node state conversion and causal path optimization mechanism, each node v in the network i It can be updated at every moment based on historical data, causal paths, and timing information, and reflects the evolution of each node in the entire system. Ultimately, by comprehensively reasoning about the status of all nodes, the prediction results of each decision point in the insurance data can be obtained; The final prediction result can be expressed by the following formula: Among them, H i Represents node v i The final state, P(Y i |H i ) is the predicted probability based on the node state. T represents the time step used for prediction reference; A125.5 is accelerated by hardware specifically designed for graph network-enhanced insurance data audit optimization methods; Due to the excessive and limited use of symbols, the mathematical symbols used in A12 are only represented in A12 and have no relationship with the symbols with the same representation described in other steps; A13. In addition to auditing the data of the entire process of calculation-prediction-analysis by means of graph modeling, the present invention also provides a virtual-real dynamic deductive data audit optimization method. As a collaboration, this method performs virtual simulation modeling on actual data, and uses the roof leakage of houses, the years of construction of houses in the block, and the historical accident data of the same block as references to perform dynamic deductive reasoning, and audits the data by means of counterfactual reasoning, as follows: A131. First, through virtual simulation modeling of actual data, a dynamic deductive model corresponding to the actual data is constructed, and factors such as house roof leakage, construction age of houses in the block, and historical accidents are set as input data. These data will be used as the initial conditions of the virtual simulation model, and the evolution process of the house at different time nodes and under different conditions will be generated through deductive reasoning; Specifically, the construction of the virtual simulation model can be expressed as: X t =f(X t-1 ,C t ,E t ) Among them, X t represents the virtual data state at time step t, C t is the external conditions at time t (such as neighborhood environment, house age, etc.), E t is the error term or random disturbance in the deductive reasoning process. Function f represents the dynamic deductive process, which describes the law of virtual data state changing over time. It is fitted using periodic functions. The dynamic deductive process is based on the theory of contradiction and Socratic syllogism. A132. Based on virtual simulation modeling, the counterfactual reasoning method is introduced to audit the data. It can evaluate the possible results in different scenarios by simulating different causal paths in the absence of a clear causal chain, and avoid Simpson's paradox. Set the counterfactual reasoning model as: in, represents the counterfactual prediction result of node i, f cf It is a counterfactual reasoning model that uses virtual simulation modeling and dynamic deductive reasoning to generate conditions and environments to derive results under hypothetical situations; Using counterfactual reasoning, we can simulate data changes under different scenarios; A133, Data Audit Optimization The optimization process of data auditing consists of the following steps: A133.

1. Data collection and preprocessing: Collect real-world observational data containing multi-dimensional data, standardize and clean it, and ensure the reliability and integrity of data quality; A133.

2. Virtual simulation modeling: Based on the collected actual data, a virtual simulation model is constructed. The model maps the actual data through dynamic simulation and virtual environment to support further reasoning and analysis of the data; A133.

3. Counterfactual reasoning: Use counterfactual reasoning to deduce virtual data and speculate how the data will change if certain conditions change. Specifically, the process of counterfactual reasoning can be represented by the following general mathematical function: Let X i is the actual observation data of data point i, which includes multiple variables (including the age of the house and environmental impact), C i are external environmental factors (including regional policies and historical risks), while A i is an assumption (for example, changing the value of a variable), the counterfactual reasoning model can be described by the following function: in, represents the counterfactual prediction result of data point i under given virtual assumptions, f cf is the counterfactual inference function, which combines the actual observation data X i 、External environment C i , error term E i and assumed change factor A i , generate new prediction values; For example, suppose X i contains historical performance data of a system, and A i is a change in an operating parameter of the system (including changing the parameter value A param ), the impact of the parameter change on system performance (including failure rate and efficiency) can be inferred. At this time, the counterfactual reasoning process is described as: By changing the assumption A param , can generate different counterfactual outcomes to evaluate system performance under different assumptions; A133.

4. Audit and Optimization: Based on the results of counterfactual reasoning, conduct data audits to identify potential errors, defects, and inconsistencies in the data. Based on the audit results, further optimize the data model and reasoning algorithm to improve the accuracy and consistency of the prediction results. A134. Mapping from virtual environment to real business scenario A134.1 Virtual Data Generation: In the previous step, simulation data has been generated by the virtual simulation model. The data is obtained by simulating the real data in a virtual environment. The simulation data can be expressed as: Among them, g sim is the simulation generation function, which is based on the input real data X i , external environment C i And the assumption A i Generate output data in a virtual environment; A134.

2. Data mapping optimization goal: In order to achieve the mapping of simulation data to real data, it is necessary to construct a mapping function M sim2real , through this function the simulation data Convert to real data Y i real , the function can be expressed as: Among them, B i is an additional correction parameter in actual observation, which represents the conversion error or deviation from the simulation model to the real world, C i Representing the external environment, the mapping function makes the simulation data consistent with the real-world data in terms of physical quantity, spatiotemporal relationship and other relevant data characteristics through reverse optimization and correction of the correction factor; A134.

3. Optimization algorithm and mapping accuracy improvement: In order to improve the mapping accuracy, we use real-world data Y i real To optimize the mapping function M sim2real ; The mapping accuracy is optimized by minimizing the following objective function: The minimization process of this objective function can be achieved by the gradient descent method, and the parameters of the mapping function are adjusted by back-propagating the error; A134.

5. Reverse correction and modification: In the process of converting simulation data into real data, it is necessary to consider the influence of the external environment and other uncertain factors. By introducing a reverse correction mechanism, the accuracy of the mapping can be further optimized, including using historical data or expert knowledge to adjust the parameter B in the mapping function. i , making the model more adaptable to complex factors in the real environment; A134.

6. Final data audit and optimization: Finally, based on the optimized data mapping function M sim2real , accurately connect simulation data with real data, and through the data audit process, check data consistency and validity, identify potential anomalies or errors, thereby optimizing data quality and improving the accuracy of predictions and decisions.

3. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The hardware acceleration method B0 of the insurance data audit optimization method specifically for graph network enhancement includes: Hardware components: B1. Dedicated Processing Unit (PU): This hardware accelerator design includes multiple dedicated processing units, which are used for the creation and update of graph data structures, enhanced computing of graph networks, and deep audit optimization of graph data. Each PU can process different graph nodes and edge information in parallel, supporting efficient graph traversal, shortest path calculation, and connectivity analysis operations. B2. Graphics Processing Unit (GPU): With the integrated graphics processing unit (GPU), the hardware can efficiently perform training and inference calculations of parallel graph networks. GPU supports graph modeling and optimization of large-scale insurance data by implementing parallel computing at the hardware level, significantly reducing the bottleneck of traditional CPU computing. B3. Graph Storage Unit (GSU): To support the storage and reading of large-scale insurance data graphs, the present invention designs a special graph storage unit (GSU), which adopts a high-bandwidth storage architecture, supports large-capacity graph structure data storage, and can perform fast data exchange between the memory and the processing unit, ensuring efficient access to graph data; B4. Hardware Acceleration Interface (HAI): The hardware accelerator is designed with a dedicated hardware acceleration interface for data transmission and interaction with external systems or data sources. Through the HAI interface, insurance data can be quickly transferred to the hardware accelerator for processing, and the processing results can also be returned in real time, supporting real-time data auditing; Hardware acceleration process: B5. Insurance data input: External insurance data is input into the hardware accelerator through the HAI interface. The input data includes historical insurance claims records, customer information, and insured product data. B6. Graph network construction: The dedicated PU receives the input insurance data and constructs an insurance data graph, which contains nodes (including customers, policies, and claims events) and edges (including associations and claims history). This process is executed in parallel by the graphics processing unit in the hardware accelerator, which can quickly complete the construction of large-scale data graphs; B7, graph enhancement and optimization: After the graph data is constructed, the GPU performs deep learning training and audit optimization on the graph based on the module A0. The optimization tasks include predicting potential risks, analyzing customer behavior, and identifying possible abnormal events. B8. Output audit results: The processed audit results are transmitted back to the external system through the HAI interface to ensure the real-time and accuracy of the audit results.

4. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The hardware acceleration method C0 dedicated to the dynamic deductive data audit optimization method combining virtuality and reality includes: Hardware components: C1. Simulation Processing Unit (SPU): The simulation processing unit is dedicated to the generation and processing of virtual simulation data. Driven by the real-world business data simulation model, the unit can generate simulation data based on the input real data and support dynamic parameter adjustment and environmental simulation. SPU can simulate different simulation scenarios according to real-time input conditions; C2. Data Mapping Accelerator (DMA): The data mapping accelerator is responsible for mapping between simulation data and real data. The accelerator uses a customized mapping algorithm (including reverse optimization and correction mechanisms) to achieve efficient conversion from virtual data to real data. The DMA accelerator supports large-scale data mapping and can quickly respond to real-time data input. C3, Dynamic Fusion Module (DFM): This module is used to process the fusion of virtual and real data. DFM can dynamically combine virtual simulation data with real-world data, automatically adjust simulation parameters to adapt to changes in the real environment, and improve the accuracy of data fusion through optimization algorithms; C4, Multi-dimensional Processing Compute Unit (MPCU): This computing unit is designed to efficiently perform complex multi-dimensional dynamic deductive computing tasks. By combining hardware-level parallel computing architecture, MPCU can process multiple types of data simultaneously and support parallel processing of spatiotemporal data, dynamic deductive reasoning results, and counterfactual reasoning; C5, Hardware Collaborative Interface (HCI): This interface is used to exchange data with external hardware or cloud computing platforms. The HCI interface supports high-speed data transmission, ensuring that data can be seamlessly transferred from the hardware acceleration platform to the external system for further analysis and decision support; Hardware acceleration process: C6. Real data input: External real data is input into the hardware platform through the HCI interface. The data content includes real-time monitoring data and environmental data. C7, Virtual simulation generation: The simulation processing unit (SPU) generates simulation data in the virtual environment based on the input real data. The simulation model generates the expected data scenario based on the physical model, historical data and external variables; C8, Simulation data and real data mapping: Through the data mapping accelerator (DMA), the system connects the simulation data with the real data to ensure that the simulation results can accurately reflect the dynamic changes in the real world; C9, Dynamic deduction and optimization: The dynamic data fusion module (DFM) performs dynamic deductive reasoning on the data through reverse reasoning and correction mechanisms. The optimization process includes compensation for uncertainty factors and adjustment of the deductive model; C10. Audit result output: The final audit result is output to the external system through the HCI interface for generating decision support reports or real-time monitoring.

5. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The audit module D0 constructed based on the architecture B0 and the architecture C0 described in the method B0 and the method C0 includes: D1. Processing Unit (PU): This unit is responsible for executing the computational tasks of graph network enhancement and virtual-real dynamic deductive data audit. The processing unit includes multiple dedicated computing cores for parallel processing of complex graph data structures and simulation calculations, ensuring efficient data processing and reasoning optimization during the audit process. D2, Memory Unit: The memory is used to store intermediate calculation results related to graph data, historical audit data, virtual simulation models, and counterfactual reasoning results. It supports high-speed read and write operations through large-capacity and high-bandwidth storage to ensure that there is no bottleneck when executing complex computing tasks. D3. Simulation Module: This module is specifically used to generate and process virtual simulation data according to the virtual-real combination solution in Method C. With the support of hardware acceleration, the simulation module can generate virtual scenes based on real data in real time to simulate the audit process in different environments. D4, Data Mapping Accelerator (DMA): This component is responsible for performing the mapping calculation between virtual simulation data and real data, ensuring seamless data conversion from the simulation environment to the real environment, and ensuring the authenticity and accuracy of the final audit results; D5, Interface Module: It is used for data interaction between the audit server and the external system or network environment. The interface module supports high-speed data transmission and instruction execution, ensuring that the input full business chain process data can be quickly transmitted to the audit server and quickly returned to the external system after the audit results are generated; D6, Graph Data Processing Unit: This unit is specifically used to process graph network-related tasks, including graph data construction, node and edge processing, and graph enhancement and optimization. Through hardware-level parallel computing, the Graph Data Processing Unit can significantly improve the efficiency and accuracy of graph data operations. D7. Dynamic Inference Module: The dynamic inference module is responsible for performing counterfactual reasoning and dynamic deductive analysis based on the graph network and dynamic deductive algorithm in method B. The memory D2 of this module stores a computer program D1. When the computer program D1 is executed, the module performs an audit operation on the entire business chain process data according to method A. The computer program D8 contains an instruction set for coordinating the collaborative work of the above-mentioned hardware components to realize the full-process audit of data input, processing, simulation, optimization and output. When the program D8 is executed, the audit server first receives the external business chain process data and transmits it to each processing unit through the data interface module. The graph data processing unit constructs and optimizes the graph network model. The simulation module generates a virtual scene based on real-time data and realizes the data conversion of virtual and real through the data mapping accelerator. Finally, the dynamic reasoning module performs in-depth audit analysis on the data through counterfactual reasoning and outputs an audit report.

6. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The method E0 for distributed storage of insurance data after graph modeling includes: E1. Data Sharding: When performing graph modeling, the insurance data is first sharded according to certain partitioning rules. Each piece of data contains a certain range of node and edge information to ensure that the data is evenly distributed among different computing nodes. Assume that the graph data is in is the node set, and E is the edge set. The data sharding process can be expressed as: in, is the i-th data shard, N is the total number of shards, and each Contains node sets and edge set E i The sharding rules can be determined based on node degree, graph topology, or edge weights; E2. Node-to-Compute Mapping: Store the sharded data on different compute nodes to ensure that each compute node can store the corresponding data fragments and can perform the graph data computing tasks nearby. Assume there are M compute nodes. Each computing node Storage Sharding The mapping relationship can be expressed as: This mapping is dynamically adjusted based on load balancing and data dependencies to ensure that the computational tasks processed by each computing node minimize cross-node communication; E3. Compute-Storage Co-location: Each computing node not only stores the shard data, but also contains the related computing units, including GPU and CPU. The computing task of each computing node can be expressed as: Among them, f(v) represents the The calculations performed (such as node feature extraction) are as follows: g(e) represents the calculation of edge e∈E i The calculations performed (such as edge weight calculations) are limited to the locally stored data at each node, ensuring that the calculation process does not involve cross-node data access, thereby improving calculation efficiency; E4. Data Synchronization & Consistency: Since data is stored on different computing nodes, this method introduces an efficient distributed synchronization mechanism to ensure the consistency of data among all computing nodes. The data synchronization process is implemented through a consistency protocol, such as Paxos or Raft protocol. Assume that Modified the data The synchronization process can be expressed as: in, Is a node The data on is the modified increment, It is a data merging operation. After synchronization, the data of all computing nodes will be updated to a consistent state; E5. Data query and distributed query engine: In order to support efficient data query in a distributed storage environment, this method designs a distributed query engine. The query request of the query engine can be expressed as: Among them, Q is the query operation, It is graph data. are query parameters (including path query and node feature query), It is the query result. During the query process, the query engine automatically determines the storage location of the query data and distributes the request to the corresponding computing node for processing. The cross-node query is merged through: in, is the computing node C i The query results returned are: is the final query result; E6. Fault Tolerance: In a distributed storage system, in order to ensure that the system can still work normally when some computing nodes fail or crash, this method introduces a fault tolerance mechanism. Data backup can be achieved through a copy strategy. Assuming that the data There are K replicas on multiple computing nodes, which can be expressed as: in, Is a node The kth copy on the node If a failure occurs, the system will automatically switch to the data on that replica for calculation to ensure that the task can continue to execute.

7. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The insurance data processing method F0 based on a generalized linear model (GLM) comprises: F2, data preprocessing; F21. Data cleaning, including: F21.

1. Remove duplicate data. Insurance data may come from multiple systems, such as the insurance system and the claims system, and there may be duplicate records. The removal method includes detecting the same policyholder ID, license plate number, insurance policy number and other fields. If the contents of all fields are the same, delete the duplicate items; F21.

2. Handle missing values. For continuous variables (age, credit score), use the mean, median, or mode to fill in. For categorical variables (car brand, agent ID), use the mode (most common category) to fill in or set it to the "unknown" category. F21.3 detects and processes outliers, calculates the interquartile range (IQR), removes values ​​outside the range of [Q1-1.5IQR, Q3+1.5IQR], and if the data follows a normal distribution, removes values ​​exceeding the mean ± 3 times the standard deviation, and uses Winsorization to truncate outliers and return them to a reasonable range; F22, feature screening, including: F22.1, Correlation analysis, calculate the correlation between the feature variables and the target variable, and select the variables with stronger correlation, including: if it is found that the vehicle color has no obvious effect on the claim amount, then the variable can be deleted; F22.2, variance filtering, calculate the variance of each feature variable and eliminate variables with too low variance (i.e., almost all samples have the same value); F22.3, Recursive Feature Elimination (RFE), train a preliminary GLM model, calculate the regression coefficient (β) of each variable, and gradually remove features that contribute less to the prediction; F22.

4. Combined with the knowledge of business experts, in some cases, relying solely on statistical methods to screen features may not be enough. It is necessary to combine industry experience and use the intelligent insurance data processing and decision optimization method based on a large language model and human expert feedback H0; F22.

5. The age of the insured affects the claim rate. Past claims records are crucial for fraud detection. Social network relationships reflect potential group fraud. F23, variable conversion, including: F23.1, categorical variable coding, including: One-Hot Encoding: suitable for unordered categories, including "car brand" (Toyota, Honda, BMW); Label Encoding: Applicable to ordered categories, including "credit rating" (A, B, C); Target Encoding: Calculate the target mean for each category, including the impact of “insured region” on the claim rate; F23.2, numerical variable transformation, including: Log Transformation: used for right-skewed distribution data, for variable x, including the amount of compensation: x ′ =log(1+x) Square Root Transformation, for variable x, is suitable for slightly right-skewed data: Box-Cox transformation x is used to make the data close to normal distribution and improve model stability; F23.

3. Standardization, including: Z-score standardization; Among them, μ is the mean and σ is the standard deviation; Min-Max normalization; Variables that apply to a specific range, in the insurance field include credit scores (between 300-900); Robust Scaling Among them, Q1 and Q3 are the first and third quartiles, respectively, and are applicable to variables with extreme values, including claim rates; F23.

4. Handling data imbalance, including: Undersampling: Randomly remove some majority class samples to make their ratio with the minority class more balanced. The undersampling method reduces the number of majority class samples by randomly removing majority class samples, thereby achieving class balance of data; Oversampling: Oversampling methods balance the dataset by synthesizing samples of minority categories. Using the SMOTE (Synthetic Minority Over-sampling Technique) method, new synthetic samples can be generated between minority category samples through interpolation technology; Weighted loss function: During the GLM training process, a higher weight is given to the minority class to reduce the impact of class imbalance on the model. During the training process, a higher weight is given to the minority class through the weighted loss function to reduce the impact of class imbalance on the model. This method can make the model more sensitive to minority class data during the training phase and improve its classification accuracy; F23.5, data division, including: Training Set (70%): used for model training; Validation Set (15%): used to adjust hyperparameters; Test Set (15%): used for final evaluation; The division method is Out-of-time Test: the future data set of the training set is used for testing; F3. Select the GLM model, including: Poisson Regression: Applicable to predicting claim frequency, i.e. the number of times a claim occurs. The Poisson regression model can effectively describe the distribution of the number of times an event occurs. It is often used to analyze scenarios where the number of claims is small but occurs frequently. Gamma Regression: Applicable to predicting the claim amount, that is, the average cost of each claim. Gamma regression is suitable for modeling continuous data with positive skewed distribution. Log-Normal Regression: It is suitable for predicting claim amounts with large fluctuations. The log-normal regression model can handle claim amounts with large fluctuations or long-tail distributions. Tweedie Regression: Suitable for modeling mixed data of zero claims and positive claims. Tweedie regression can handle data containing a large number of zero values ​​(no claims) and a small number of positive values ​​(claims), which is very suitable for the zero artifact problem in insurance claim prediction. F4. Model training, including: F4.

1. Define the GLM model: Where g(·) is the link function, X is the input variable, Y is the target variable, and β is the parameter to be estimated; F4.2, Maximum likelihood estimation solution parameters: in: The parameter value obtained by maximum likelihood estimation, that is, the best parameter estimate obtained by optimizing the likelihood function, β: the parameter to be estimated of the model, usually the regression coefficient or other model parameter, P(Y i |X i ,β): Given an input variable X i and parameter β, the target variable Y i The conditional probability density function or probability mass function, logP(Y i |X i ,β): Taking the logarithm of the likelihood function (i.e., log-likelihood function), which helps to simplify the calculation and enhance numerical stability, N: number of samples, indicating the total number of samples in the data set; F4.

3. Training process, including: Gradient descent is used to optimize parameters. In the insurance field, gradient descent can effectively optimize the regression coefficients for predicting insurance claim frequency or claim amount; Combined with cross validation, model hyperparameters and regularization parameters are optimized. It evaluates the generalization ability of the model by dividing the dataset into multiple subsets (usually k subsets) and training and testing the model multiple times. In each training process, the selection of hyperparameters is optimized through cross validation. Calculate AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) to select the optimal model. AIC and BIC evaluate the model by balancing the fit and complexity of the model, avoiding overfitting (the model is too complex) or underfitting (the model is too simple). In the insurance claim prediction, the AIC and BIC values ​​of different models (such as Poisson regression, gamma regression, etc.) can be compared to select the model that can best predict the insurance claims. F4.

4. Prediction optimization, including: Bias Correction: Adjust model bias based on residual analysis to improve prediction accuracy. In the insurance field, residual analysis helps identify model errors in specific customer groups or specific types of claims, thereby adjusting the model. Risk Segmentation: Divide customer risk levels based on prediction results. By analyzing customers’ historical claim data, personal information, and predicted claim probability, detailed risk segmentation can be performed on customers, helping insurance companies determine the risk levels of different customer groups. Dynamic Pricing: Adjust personalized premiums based on risk level. As the customer's risk level changes, the insurance company can adjust the premium in real time to ensure the fairness of the premium and the financial stability of the insurance company.

8. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The information disclosure and data security control method G0 designed for the insurance industry includes: G1. Data Encryption Storage: To ensure that data will not be illegally accessed or leaked during storage, all sensitive insurance data must be encrypted and stored. Suppose the insurance data set is Each data item d i It may contain sensitive content such as user's personal information, insurance contract terms, and claims records. The data encryption process can be expressed as: Among them, Enc(d i ,k) is the encryption function, k is the encryption key, It is an encrypted data item. The encrypted data is stored in the insurance company's data center and is provided to the outside through authorized access. G2, Privacy Protection & Anonymization: In order to meet the requirements of privacy protection regulations, especially when it comes to personal information processing, it is assumed that the data Contains personally identifiable information (PII), and uses data de-identification technology (k-anonymity) to replace or encrypt sensitive information to remove identifying marks, making the data Transformed to a dataset without personally identifiable information Data anonymization ensures that personal data is not leaked, so that data analysis and processing can still be carried out under the premise of complying with privacy protection; G3, Government Interface & Transparent Data Disclosure: Allow government regulators to access encrypted stored data through authentication and conduct necessary data audits and monitoring. The implementation of this interface provides the government with a set of controlled access rights and operation permissions, enabling the government to access specific model prediction results and data without leaking sensitive information of the enterprise. Assume that the government interface is The interface access process can be expressed as: in, Represents the government access interface. Authenticate(gov_key) is the government authentication process. gov_key is the key or certificate provided by the government. A transparent view of data accessible to governments that contains only non-sensitive information, enabling governments to audit and monitor the insurance industry’s use of data while protecting personal and business confidential information; G4. Prediction Transparency: By making part of the prediction model process public and transparent, the government can audit data processing and model reasoning while protecting the trade secrets of insurance companies. Specifically, the intermediate results of model training and prediction (model parameters and predicted decision boundaries) can be displayed to the government through encryption or "explainability" technology (such as SHAP or LIME): in, It is an insurance dataset and the prediction process under model parameters θ, The forecast results are viewable by the government. In this way, insurance companies can provide sufficient information to the government without revealing model parameters to ensure the fairness and transparency of the forecast process; G5. Data Monitoring & Auditing: Record all access requests, data modifications, and prediction operations. All audit logs will be encrypted and stored in a dedicated audit database for subsequent review and analysis. The recording formula for the audit log is: in, It is an audit log recorded at time point t. action indicates the type of operation (data access, prediction request, and execution method A), timestamp is the time when the operation occurred, user_id is the user ID who performed the operation, and data_accessed is the data item accessed. These logs provide traceable audit records for governments and insurance companies to ensure the legitimacy of data use. Disaster Recovery & Data Recovery: In order to cope with data loss or tampering caused by emergencies, all data copies will be backed up regularly during the storage of insurance data, and the backup data will be stored in different geographically dispersed data centers. The data recovery process can be expressed as: D recovered =Recover(D,backup_location) Among them, D recovered is the recovered data, D is the original data, and backup_location is the location where the backup data is stored. Through this mechanism, when hardware failure or data loss occurs, the insurance company can quickly restore the data and continue to ensure the normal operation of the business.

9. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The intelligent insurance data processing and decision optimization method H0 based on a large language model and human expert feedback includes: H1. Preprocessing of multi-source heterogeneous data: The data comes from structured data (such as customer information, historical claims records) and unstructured data (customer comments, social media information) from insurance company databases. In order to process these data, the present invention uses a large language model to clean, denoise, fill missing values, and convert the text data into a structured format to ensure data consistency and integrity: D processed =f(D raw ,LLM) Among them, D raw represents the original data set, D processed represents the processed data set, f is the data processing function, and LLM represents the large language model; The processed data is used as the basic data of the production process of the present invention and is used in all processes of method F; Intelligent decision support during model training: During model training, the large language model is used to provide intelligent decision support, including: Dynamic hyperparameter adjustment: Large language models automatically optimize hyperparameters (such as learning rate, batch size, etc.) based on real-time feedback during training, accelerating training and improving model performance; Business logic optimization: The large language model is combined with current business needs to optimize the business logic of the model to ensure that the model output complies with insurance business rules; R=LLM(F,D processed ) Among them, F is the feedback during the training process, and R is the optimization strategy provided by the large language model; For the optimized and iterated models, we continuously select the best version and use it as the basis for preprocessing the aforementioned multi-source heterogeneous data. H2. Prediction evaluation and optimization feedback: In the prediction phase, the present invention uses a large language model and human expert feedback to optimize the prediction results to ensure the long-term effectiveness of the model, including: Prediction result analysis: The large language model can analyze the results of model prediction, identify possible errors, and adjust the prediction results; Real-time feedback and adjustment: Based on the feedback of real-time data and historical data, the large language model continuously optimizes the prediction strategy to improve the prediction accuracy; Business decision support: Based on the optimized prediction results, the big language model provides strategic suggestions to decision makers to ensure the scientificity and feasibility of the decision; in, is the original prediction result, is the optimized prediction result. The prediction results are used as a partial reference and comprehensively analyzed with the feedback information provided by human experts to evaluate the results of data application such as optimized predictions.

10. The distributed clustering model training device based on insurance big data and its application as claimed in claim 1, characterized in that: The low-cost distributed cluster architecture 10 includes: I1. Hardware components, including: Audit module D0, independent of other modules, independently audits all business data in the storage server in the cluster and provides audit reports on a periodic basis; The computing server adopts an x86 architecture, including: a graphics card, a CPU, a memory, an RJ45 network card, a QSFP+ network card, an IB network card, a hard disk, and a computer program ICODE. When the computer program ICODE is executed by the CPU, the computing server implements the method F0 described in F; Network switches, including: RJ45, QSFP+, IB switches; The storage server adopts an x86 architecture, including: a CPU, a memory, an RJ45 network card, a QSFP+ network card, an IB network card, an array card, a mechanical hard disk, and a computer program ISTORE. When the computer program ISTORE is executed by the CPU, the computing server implements the method E0 described in E; A data exchange server storing a computer program G1, which, when executed, causes a processor to execute the information disclosure and data security control method G0 designed for the insurance industry; Infrastructure, including: external three-phase power supply, monitor, input and output devices, external network; I2. Network, including: RJ45 network: Enable DHCP, access the external network, take into account the BMC ports of each device, and set up a gateway; QSFP+ network: DHCP is not enabled, no external network access, and no gateway is set; IB network: DHCP is not enabled, no external network access, no gateway; I3. Storage architecture, including: Storage and computing integration: The SSD hard disks carried by each computing server jointly form a distributed storage system. The system uses method E0 to balance the network load while improving the computing read and write IO efficiency; The storage server regularly performs complete offline backup of business data, and the backup methods include full backup and incremental backup; I4. User interaction, including: The method F comprises: a data uploading module, a data preprocessing module, a feature engineering module, a training and evaluation module, a prediction and optimization module, and an interactive feedback and reporting module; Method G includes: login and authentication interface, data encryption storage settings, data privacy protection and anonymization, government interface and data disclosure management, forecast transparency review, audit and data monitoring logs, disaster recovery settings and data backup, and data security policy and compliance checks; The method H includes: a data preprocessing interface, an intelligent decision support interface, a prediction evaluation and optimization feedback interface, a comprehensive feedback and expert advice interface, and a model training and prediction history record.

Citation Information

Patent Citations

  • Defect high-risk module identification method based on software network

    CN110147321A

  • Data distributed storage management system oriented to edge device

    CN115733848A

  • Insurance anti-fraud prediction method and system based on machine learning

    CN116911882A

  • Data analysis model training method and device, equipment, storage medium and product

    CN116956005A

  • Risk estimation method and system based on scheduling operation behavior deduction

    CN118572657A

Cited By

  • Account risk prediction method and system based on machine learning

    CN120450708A

  • A machine learning based account risk prediction method and system

    CN120450708B

  • Adaptive multi-strategy fusion big language model training optimization method

    CN121615726A

  • An adaptive multi-strategy fusion large language model training optimization method

    CN121615726B