Giant disaster reinsurance contract optimization method, system, equipment and medium
Optimizing catastrophe reinsurance contracts through multi-agent reinforcement learning and game mechanisms has solved the lag of traditional methods and the neglect of multi-party cooperation, realized the dynamic adjustment of contract terms and automation of risk management, and improved decision-making efficiency and system stability.
Patent Information
- Application Number
- CN202510395360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The traditional catastrophe reinsurance contract evaluation method has lag, lacks a dynamic adjustment mechanism, and is unable to effectively respond to the risks and challenges brought about by climate change and frequent disasters. It also ignores cooperation and game between multiple participants, resulting in inefficient decision-making.
The multi-agent reinforcement learning method is adopted, combined with the game mechanism, and a dynamically optimized reinsurance contract mechanism is designed. Through the collaboration and game between insurance companies, reinsurance companies, risk assessments, regulatory compliance and external environmental agents, the contract terms are adjusted in real time, and parameters such as reinsurance ratio, reinsurance rate and deductible are optimized.
It has achieved dynamic optimization of the terms of the reinsurance contract, and can automatically adjust in market changes and disaster events, ensuring the balance of risk sharing and profitability, improving decision-making efficiency and system stability, and reducing artificial deviations and risk management costs.
Smart Images

Figure CN120338964A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of insurance, and particularly relates to an optimization method, system, device and medium for catastrophe reinsurance contracts. Background Art
[0002] Catastrophe reinsurance is an insurance product that provides risk sharing for insurance companies, mainly dealing with huge economic losses brought by natural disasters such as earthquakes, typhoons, floods or other large-scale emergencies. Traditional catastrophe reinsurance contracts are usually negotiated and formulated between insurance companies and reinsurers based on historical data, statistical models and manual pricing strategies. However, with the global climate change and frequent disasters, traditional catastrophe reinsurance models face many new challenges and deficiencies.
[0003] Traditional catastrophe reinsurance contracts usually conduct risk assessment based on historical data and disaster models. However, this method relies on historical disaster events and ignores emerging and unappeared risk factors. With the continuous changes in climate and social environment, past risk patterns may not accurately predict future disaster risks. Therefore, traditional risk assessment methods have lag, and it is difficult to effectively respond to new risk challenges. Currently, the terms of catastrophe reinsurance contracts are usually set once, lacking sufficient flexibility and dynamic adjustment mechanisms. With the changes in market environment, risk exposure and disaster occurrence frequency, the contract terms may need to be adjusted, but traditional methods cannot quickly respond to these changes, resulting in the contract terms may not match the market demand, thus affecting the effectiveness and fairness of the contract. Moreover, the cooperation and game among multiple participants are ignored, leading to low efficiency in the decision-making process. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an optimization method, system, device and medium for catastrophe reinsurance contracts, which are used to wholly or at least partially solve the technical problems existing in the above-mentioned prior art, such as the lag in the assessment method, the lack of dynamic adjustment mechanism, and the ignoring of the cooperation and game among multiple participants.
[0005] In a first aspect, an embodiment of the present application provides an optimization method for catastrophe reinsurance contracts, including: Constructing multi-agents; Adopting a mathematical optimization model, combined with the core terms in the reinsurance contract, to design a dynamically optimized reinsurance contract mechanism to optimize the reinsurance contract terms, wherein optimizing the reinsurance contract mechanism includes a reinsurance contract parameter optimization strategy and a multi-agent collaborative optimization of contract parameters; Design the collaboration and game mechanisms among the agents, and adopt the multi-agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real time. Among them, the collaboration and game mechanisms among the agents include the game mechanism between the insurance company and the reinsurer, and the cooperation mechanism of the agents.
[0006] Optionally, the constructed multi-agents at least include an insurance company agent, a reinsurer agent, a risk assessment agent, a regulatory compliance agent, and an external environment agent.
[0007] Optionally, the reinsurance contract parameter optimization strategy includes: cession ratio optimization, reinsurance rate optimization, and deductible optimization; Among them, the cession ratio is optimized according to the following formula:
[0008] In the formula, is the optimal cession ratio, represents the penalty term when the capital adequacy ratio does not meet the regulatory requirements, is the penalty coefficient, weighing the profit and the SCR violation risk, is the expected profit of the product line.
[0009] The reinsurance rate is optimized according to the following formula:
[0010] In the formula, P* is the optimal reinsurance rate, E[L|L>d] is the expected loss after exceeding the deductible d, and ζ is the profit loading factor of the reinsurer.
[0011] The deductible is optimized according to the following formula:
[0012] In the formula, is the optimal deductible, is the degree of risk exposure, is the risk adjustment parameter, is the expected profit of the product line.
[0013] Optionally, the multi-agent collaborative optimization of contract parameters includes: For the insurance company agent, adjust the cession ratio based on the capital adequacy ratio and market conditions; For the reinsurer agent, adjust the reinsurance rate and deductible according to the market competition situation; For the regulatory compliance agent, provide real-time compliance checks to ensure that the capital adequacy ratio meets the regulatory requirements; For the risk assessment agent, provide market risk forecasts to influence the decisions on the cession ratio and deductible; For external environmental agents, monitor the market economy situation and affect premium rate adjustment.
[0014] Optionally, the game mechanism between the insurance company and the reinsurer includes: Under stable market conditions, the insurance company and the reinsurer optimize the contract through Stackelberg reinforcement learning. The reinsurer first formulates the reinsurance contract parameters, and its goal is to maximize profits. Among them, the reinsurance contract parameters include the reinsurance premium rate, the ceding percentage, and the deductible; After the reinsurer sets the contract parameters, the insurance company selects the optimal reinsurance purchase strategy based on its own capital adequacy ratio and market risks; When the market fluctuates violently or a disaster event occurs, the insurance company and the reinsurer choose to cooperate and share the benefits according to the Shapley value allocation mechanism.
[0015] Optionally, the multi-agent reinforcement learning method includes: When each agent makes a decision, it defines a state vector for the market environment, contract terms, and historical data; For the insurance company agent, set the ceding percentage and whether to accept the reinsurance contract; for the reinsurer agent, set the reinsurance premium rate and the deductible; for the risk assessment agent, whether to renegotiate the reinsurance contract terms; for the regulatory compliance agent, whether to change the minimum solvency adequacy ratio; for the external environmental agent, whether to change the risk exposure; Design the agent reward function and adopt the Actor-Critic architecture to optimize the contract parameters. Among them, use the Actor network to generate the optimal contract strategy, use the Critic network to evaluate the performance of the current strategy in the long-term benefits, and adjust the strategy; Initialize the neural network of each agent, the ReplayBuffer for storing the historical interaction data of the agent, and set the learning rate and discount factor; Each agent interacts with the environment based on the current strategy, and obtains the state, selects actions, immediate rewards, and the next state, and stores the interaction results in the Replay Buffer; Randomly sample from the Replay Buffer for training, calculate the Q-value update of each agent, and use the deterministic policy gradient to update the agent's strategy, and use the target network for Q-value estimation; when the agent's strategy converges after multiple training cycles and the long-term return of each agent reaches the optimal, the training process terminates.
[0016] Optionally, the policy is updated according to the following formula:
[0017]
[0018] In the formula, is the learning rate, is the state-action value function, is the gradient with respect to and is the gradient with respect to .
[0019] In a second aspect, an optimization system for a catastrophe reinsurance contract provided by an embodiment of the present application includes: a construction unit configured to construct multi-agents; a design unit configured to adopt a mathematical optimization model and combine core terms in a reinsurance contract to design a dynamically optimized reinsurance contract mechanism to optimize reinsurance contract terms, wherein optimizing the reinsurance contract mechanism includes a reinsurance contract parameter optimization strategy and a multi-agent collaborative optimization of contract parameters; an optimization unit configured to design a cooperation and game mechanism between agents, and dynamically optimize game strategies by using a multi-agent reinforcement learning method to adjust contract terms in real time, wherein the cooperation and game mechanism between agents includes a game mechanism between an insurance company and a reinsurer, and a collaborative mechanism of agents.
[0020] In a third aspect, an electronic device provided by an embodiment of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the optimization method for the catastrophe reinsurance contract described above are implemented.
[0021] In a fourth aspect, a storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the steps of the optimization method for the catastrophe reinsurance contract described above are implemented.
[0022] From the above technical solutions, the following advantages of the present invention can be seen: In the optimization method, system, device, and medium for the catastrophe reinsurance contract provided by the present application, a strategy combining game and reinforcement learning is adopted to optimize the contract terms between an insurance company and a reinsurer. The game mechanism helps the agents find the optimal balance between competition and cooperation, while reinforcement learning enables the agents to continuously adjust strategies in a dynamic market. It is ensured that the reinsurance contract terms can be automatically optimized when the market changes, disaster events occur, and regulatory requirements change. The agents optimize the contract terms through a real-time feedback mechanism to ensure the balance of risk sharing and profitability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of an optimization method for a catastrophe reinsurance contract provided by an embodiment of the present invention; Figure 2 It is a construction process of a multi-agent reinforcement learning algorithm provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of an optimization system for a catastrophe reinsurance contract provided by an embodiment of the present invention; Figure 4 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0025] In the following detailed description, various embodiments of the present disclosure will be more fully described. The present disclosure can have various embodiments and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0026] In the following, the term "comprise" or "may comprise" that can be used in various embodiments of the present disclosure indicates the existence of the disclosed function or operation, and does not limit the addition of one or more functions or operations. In addition, as used in various embodiments of the present disclosure, the terms "comprise", "have" and their cognates are only intended to indicate a specific feature, number, step, operation or combination of the foregoing items, and should not be understood to first exclude the existence or addition of one or more other features, numbers, steps, operations or combinations of the foregoing items.
[0027] In various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the recited words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0028] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] Refer to Figure 1 The figure shows a flowchart of an optimization method for a catastrophe reinsurance contract in a specific embodiment, including the following execution steps: Step 100: Construct a multi-agent.
[0030] Specifically, the constructed multi-agent at least includes an insurance company agent, a reinsurer agent, a risk assessment agent, a regulatory compliance agent, and an external environment agent.
[0031] Specifically, the construction process of the multi-agent includes: For the insurance company agent, define the goals and functions: Goals: Manage insurance business, including underwriting and claims settlement, achieve profitability and maintain stable operation. Functions: Receive customer insurance applications, assess risks to determine whether to underwrite and set premiums; process claims applications, calculate and pay claims; manage insurance product inventory, and formulate insurance product strategies. Knowledge and data storage: Insurance product information library: Store detailed terms of various insurance products, insurance amount ranges, rate tables, etc. Customer information database: Record customer basic information, insurance history, claims records, etc. Risk assessment model: Used to evaluate customer risk levels and determine premium levels. Decision mechanism construction: Underwriting decision: Based on customer risk assessment results, product inventory, and market strategies, decide whether to accept insurance applications. If accepted, determine the premium. Claims decision: Calculate the claim amount and execute the claim based on insurance contract terms, claim application materials, and investigation results. Communication interface design: Interface with customers: Receive insurance applications and claims requests, and feedback processing results. Interface with reinsurers: Negotiate reinsurance business, such as reinsurance ratios, premium payments, etc. Interface with the risk assessment agent: Obtain customer risk assessment reports. Interface with the regulatory compliance agent: Receive regulatory update notifications to ensure business compliance.
[0032] Define the objectives and functions for the reinsurer agent: Objectives: By assuming part of the risks of insurance companies, obtain profits and control its own risk exposure. Functions: Evaluate the reinsurance needs of insurance companies, determine reinsurance terms and rates; Manage the reinsurance business portfolio for risk diversification and optimization; Provide reinsurance claims support. Knowledge and data storage: Reinsurance product library: Contains the terms of various reinsurance products, reinsurance proportion ranges, rate structures, etc. Insurance company information library: Records the business scale, risk status, historical reinsurance records, etc. of partner insurance companies. Risk assessment model: Used to evaluate the risks transferred by insurance companies. Decision-making mechanism construction: Reinsurance decision-making: Based on the risk assessment results of insurance companies, its own risk tolerance, and market conditions, decide whether to accept reinsurance requests and determine reinsurance terms and rates. Claims decision-making: During reinsurance claims, calculate the indemnity amount and make payments according to the terms of the reinsurance contract and the claims situation of insurance companies. Communication interface design: Interface with insurance companies: Receive reinsurance applications, feedback reinsurance decision results, process claims notifications and settlements. Interface with external environment agents: Obtain macroeconomic data, industry risk indices, etc. to assist in decision-making.
[0033] Define the objectives and functions for the risk assessment agent: Objectives: Provide accurate risk assessment services for insurance companies and reinsurers. Functions: Collect and analyze various types of risk-related data, and use risk assessment models to evaluate customer risks, insurance business risks, and reinsurance business risks. Knowledge and data storage: Risk data warehouse: Collects data from multiple channels, including customer information, industry data, historical claims data, market data, etc. Risk assessment model library: Stores various risk assessment models, such as credit risk assessment models, property risk assessment models, health risk assessment models, etc. Decision-making mechanism construction: Risk assessment process: According to the input data, select an appropriate risk assessment model for calculation and generate a risk assessment report. The report content includes risk levels, risk probabilities, potential losses, etc. Communication interface design: Interface with insurance companies: Receive customer insurance information and insurance business data, and feedback risk assessment reports. Interface with reinsurers: Receive reinsurance business-related data and provide risk assessment services. Interface with external environment agents: Obtain external data that may affect risk assessment, such as natural disaster data, economic data, etc.
[0034] Define the objectives and functions for the regulatory compliance agent: Objective: Ensure that the business activities of various agents in the insurance industry comply with legal and regulatory requirements. Functions: Track and interpret changes in relevant regulations and policies in the insurance industry; review whether the business processes, product terms, etc. of insurance companies and reinsurers are compliant; provide compliance training and consulting services. Knowledge and data storage: Regulatory policy database: Store domestic and foreign laws, regulations, and regulatory policy documents in the insurance industry. Compliance case database: Collect and organize compliance cases in the insurance industry for reference and learning. Decision-making mechanism construction: Compliance review process: Based on the input business data (such as insurance product terms, business operation processes, etc.), review against the regulatory policy database to determine compliance. If non-compliance issues are found, provide rectification suggestions. Communication interface design: Interface with insurance companies: Receive information such as insurance product design and business operation processes, and feedback the results of compliance reviews. Interface with reinsurers: Review the compliance of reinsurance business. Interface with external environment agents: Obtain updated regulatory policy information and update the regulatory policy database in a timely manner.
[0035] Define the objectives and functions for the external environment agent: Objective: Simulate and provide the impact of external environmental factors on the insurance industry. Functions: Collect and integrate macroeconomic data, industry dynamics information, natural disaster data, social public opinion, etc.; analyze the impact of these data on the insurance industry and provide relevant information to other agents. Knowledge and data storage: External data warehouse: Store various external data from channels such as government departments, industry associations, and news media. Environmental impact analysis model: Used to analyze the impact of external data on the insurance industry, such as the impact model of economic recession on insurance demand, the impact model of natural disasters on claim costs, etc. Decision-making mechanism construction: Data collection and analysis process: Regularly collect external data, analyze it using the environmental impact analysis model, and generate an environmental impact report. The report content includes economic trend prediction, industry risk warning, natural disaster risk assessment, etc. Communication interface design: Interface with insurance companies: Provide information such as macroeconomics and industry dynamics to help them formulate business strategies. Interface with reinsurers: Provide external environmental information that may affect reinsurance business. Interface with risk assessment agents: Provide external data to assist in risk assessment. Interface with regulatory compliance agents: Transmit updated regulatory policy information in a timely manner.
[0036] It should be noted that clarifying the roles of each agent in the multi-agent system and their optimization objectives can lay a foundation for subsequent contract optimization, game mechanism design, dynamic adjustment, and real-time feedback mechanisms. Each agent has different objectives, constraints, and interaction methods. Through the combination of multi-agent reinforcement learning (MARL) and game mechanisms, they work together to achieve final contract optimization and risk control.
[0037] Exemplarily, define the goal of the insurance company agent as optimizing the insurance company's capital adequacy ratio, reducing claim risks, decreasing reinsurance costs, and improving capital utilization efficiency; the decision variables are the cession ratio ( ), and the acceptance or rejection of the reinsurance contract (g); the constraints are to meet regulatory requirements, ensure a reasonable capital adequacy ratio, and effectively control risk exposure. Define the goal of the reinsurer agent as maximizing premium income while assuming risks, ensuring profits, and controlling risk exposure at the same time; the decision variables are the reinsurance rate ( ), and the deductible ( ); the constraints are to ensure its own capital adequacy and profitability, and at the same time provide reasonable reinsurance terms to attract insurance companies. Define the goal of the risk assessment agent as providing accurate risk prediction and assessment, assisting insurance companies and reinsurers in adjusting contract terms to ensure that risks are appropriately shared; the decision variable is whether to renegotiate the reinsurance contract terms; the constraint is to ensure the accuracy and stability of risk assessment based on historical data and market trends. Define the goal of the regulatory compliance agent as ensuring that all contract terms comply with regulatory requirements, avoiding compliance issues for insurance companies and reinsurers, and preventing default risks; the decision variable is whether to change the minimum solvency adequacy ratio; the constraint is to ensure that contract terms meet the minimum capital adequacy ratio requirements and monitor the compliance of all operations in the industry. Define the goal of the external environment agent as simulating and predicting the impact of external environments such as market changes, economic risks, and natural disasters on contract terms and agent decisions; the decision variable is whether to change the risk exposure; the constraint is to affect the capital allocation and risk management decisions of insurance companies and reinsurers, but not directly interfere with contract terms.
[0038] Step 101: Adopt a mathematical optimization model and combine the core terms in the reinsurance contract to design a dynamically optimized reinsurance contract mechanism to optimize the reinsurance contract terms.
[0039] Among them, optimizing the reinsurance contract mechanism includes reinsurance contract parameter optimization strategies and multi-agent collaborative optimization of contract parameters.
[0040] The core goal of this step is to optimize the reinsurance contract structure to ensure that the insurance company can obtain reinsurance protection at the optimal cost when the capital adequacy ratio (SCR) meets the standard, while enabling the reinsurer to obtain reasonable returns while assuming risks. For this purpose, adopt a mathematical optimization model and combine the core terms in the reinsurance contract (cession ratio, deductible, reinsurance rate, etc.) to design a dynamically optimized reinsurance contract mechanism to adapt to market changes and improve the stability and adaptability of the contract.
[0041] Exemplarily, the reinsurance contract parameter optimization strategies include: cession ratio optimization, reinsurance rate optimization, and deductible optimization.
[0042] Among them, the insurance company hopes to reduce its own risk exposure but increase the reinsurance ratio which may lead to an increase in reinsurance premium expenditure. The goal is to optimize the reinsurance ratio while ensuring compliance with the Solvency Capital Requirement (SCR), and the reinsurance ratio is optimized according to the following formula:
[0043] In the formula, is the optimal reinsurance ratio, represents the penalty term when the Solvency Capital Requirement does not meet the regulatory requirements, is the penalty coefficient, weighing the profit and the risk of SCR violation, is the expected profit of the product line.
[0044] If the market fluctuates greatly (the risk assessment agent gives an early warning), then increases, reducing the insurance company's risk exposure. If the market is stable (the claim rate decreases), then decreases, increasing the insurance company's profit.
[0045] The reinsurer needs to set the optimal premium rate p to ensure profitability while avoiding a decline in market competitiveness due to too high a premium rate. The goal is to achieve an optimal balance between the market competition environment and risk acceptability: optimize the reinsurance premium rate according to the following formula:
[0046] In the formula, P* is the optimal reinsurance premium rate, E[L|L>d] is the expected loss after exceeding the deductible d, and ζ is the profit loading factor of the reinsurer.
[0047] If the market demand increases (the insurance company increases the reinsurance demand), then increases, and the reinsurer can appropriately increase the premium rate p. If the market competition intensifies (multiple reinsurers compete for the market), then decreases, and the reinsurer reduces the premium rate to improve competitiveness.
[0048] The deductible d affects the risk sharing between the insurance company and the reinsurer, and it is necessary to ensure the optimal benefits of both. The goal is to optimize d to balance the interests of the reinsurer and the insurance company: optimize the deductible according to the following formula:
[0049] In the formula, is the optimal deductible, is the degree of risk exposure, is the risk adjustment parameter, is the expected profit of the product line.
[0050] If the risk assessment agent predicts an increase in market volatility, increase d to reduce the liability of the reinsurer for claims. If the market is stable, decrease d to increase the claims borne by the insurer and increase its profit.
[0051] Exemplarily, the multi-agent collaborative optimization of contract parameters includes: for the insurance company agent, adjusting the cession ratio based on the Solvency Capital Requirement (SCR) and market conditions ; for the reinsurer agent, adjusting the reinsurance premium rate P and deductible d according to the market competition situation; for the regulatory compliance agent, providing real-time compliance checks to ensure that the Solvency Capital Requirement (SCR) meets regulatory requirements; for the risk assessment agent, providing market risk predictions to influence the decision-making of the cession ratio and deductible d; for the external environment agent, monitoring the market economic situation to influence rate adjustments.
[0052] Step 102: Design the collaboration and game mechanisms among the agents, and adopt the multi-agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real time.
[0053] Among them, the collaboration and game mechanisms among the agents include the game mechanism between the insurance company and the reinsurer, and the collaboration mechanism of the agents.
[0054] To solve the problems of interest conflicts and information asymmetry between the insurance company and the reinsurer, in the reinsurance market, the insurance company needs to use reinsurance to reduce underwriting risks, while the reinsurer hopes to maximize its profit and control risk exposure at the same time. Therefore, the contract negotiation process can be modeled as a multi-agent game problem, involving both cooperation and competition mechanisms.
[0055] Specifically, the game mechanism between the insurance company and the reinsurer includes: Under stable market conditions, the insurance company and the reinsurer optimize the contract through Stackelberg reinforcement learning. The reinsurer first formulates the reinsurance contract parameters, and the reinsurance contract parameters include the reinsurance premium rate p, the cession ratio and deductible d, and its goal is to maximize profit:
[0056] In the formula, is the income of the reinsurer, is the premium received by the reinsurer, is the expected claim amount of the reinsurer.
[0057] After the reinsurer sets the contract parameters, the insurance company selects the optimal reinsurance purchase strategy based on its own solvency capital requirement and market risk:
[0058] In the formula, .
[0059] Insurance companies also use reinforcement learning for optimization, learning the optimal purchase strategy through deep deterministic policy gradients.
[0060] When the market fluctuates violently or a disaster event occurs, insurance companies and reinsurers choose to cooperate, optimizing contract terms through collaborative games, enabling both parties to still make a profit during market instability and ensuring market stability.
[0061] When an insurance company cooperates with a reinsurer, its total revenue is:
[0062] Insurance companies and reinsurers share the revenue according to the Shapley value allocation mechanism:
[0063] Wherein, is the revenue allocation of agent in the cooperation, is the set of agents participating in the cooperation 's total revenue.
[0064] In some embodiments, the collaborative mechanism risk assessment agent of other agents provides accurate risk assessment for insurance companies and reinsurers, especially in the case of extreme risks and disaster events, and provides policy adjustment suggestions. This agent evaluates the risks of potential extreme events (such as catastrophes, financial crises, etc.) based on historical data, market trends, and disaster models, and the reward function ( ) is:
[0065] Wherein, is the probability of the extreme event, is the expected loss caused by the extreme event.
[0066] The regulatory agency affects the contract term design of insurance companies and reinsurers through the minimum capital adequacy ratio (SCR) requirement, and the reward function ( ) is:
[0067] Wherein, is the minimum solvency adequacy ratio, is the current capital adequacy ratio, is the market adjustment factor.
[0068] The external environment agent simulates factors such as market fluctuations, economic indicators, and natural disasters, which affect the risks faced by insurance companies and reinsurers during the contract optimization process. The reward function ( ) is as follows:
[0069] where is the debt risk of the insurance company, and is the debt risk of the reinsurer.
[0070] In a specific implementation, multiple agents optimize their strategies through a game mechanism. To enable effective strategy adjustment in a complex and dynamic environment, the multi-agent reinforcement learning (MARL) method is adopted. MARL combines the self-learning ability of reinforcement learning and the theoretical framework of game theory, enabling dynamic game strategy optimization among multiple agents and real-time adjustment of contract terms according to market and risk changes to ensure the maximization of the interests of all parties. As shown in Figure 2 , the multi-agent reinforcement learning includes the following steps: S200: When each agent makes a decision, it defines a state vector for the market environment, contract terms, and historical data.
[0071] Exemplarily, the state vector is defined as follows:
[0072] where is the capital adequacy ratio of the current insurance company, is the asset-liability matching situation, is the liquidity coverage ratio, is the market volatility indicator, is the historical claim rate, is the current regulatory policy constraint, is the claim prediction value provided by the risk assessment agent.
[0073] S201: For the insurance company agent, set the reinsurance ratio and whether to accept the reinsurance contract; for the reinsurer agent, set the reinsurance rate and deductible; for the risk assessment agent, whether to renegotiate the reinsurance contract terms; for the regulatory compliance agent, whether to change the minimum solvency adequacy ratio; for the external environment agent, whether to change the risk exposure.
[0074] S202: Design the agent reward function to optimize the contract parameters.
[0075] Specifically, the global reward function is used to optimize the contract parameters while considering the goals of the insurance company, reinsurer, and regulatory agency:
[0076] Among them, is the weight coefficient.
[0077] S203: Adopt the Actor-Critic architecture. Among them, use the Actor network to generate the optimal contract strategy, use the Critic network to evaluate the performance of the current strategy in the long-term benefit, and adjust the strategy.
[0078] Specifically, the specific operation mechanism of the Actor network when generating the contract strategy is as follows: Input market information, supplier information, and platform own information into the input layer. After the information in the input layer enters the hidden layer, feature extraction will be carried out first. Each neuron will process part of the input information and extract higher-level features. For example, for market price fluctuations and supplier price historical data, the neuron can extract price trend features and judge whether the price is in an upward, downward, or stable stage. Through a non-linear activation function (such as the ReLU function) to perform non-linear transformation on the extracted features. For example, when the market demand has an obvious upward trend and the platform inventory is low, after non-linear transformation, the network can more accurately capture the information that an active procurement strategy needs to be adopted in this situation. The features extracted by different neurons are fused in the hidden layer to form a more comprehensive and representative feature vector. These fused features will be passed between the hidden layers for further processing and analysis, providing a richer information basis for the output layer to generate the contract strategy. After the processing of the hidden layer, the output layer will calculate the probability distribution of adopting different contract strategies (actions) in the current state. Suppose the contract strategy includes several dimensions such as price strategy, delivery time strategy, and quality standard strategy. For the price strategy, what may be output is the probability of choosing low price, medium price, and high price strategies in the current state; for the delivery time strategy, what may be output is the probability of choosing short delivery period, medium delivery period, and long delivery period. According to the output probability distribution, determine the final contract strategy through a certain sampling method (such as roulette wheel selection method, etc.). For example, if the probability of choosing the low price strategy is 0.6, the probability of the medium price strategy is 0.3, and the probability of the high price strategy is 0.1, then in multiple samplings, about 60% of the cases will choose the low price strategy. At the same time, combined with the strategy probabilities of other dimensions such as delivery time and quality standard, finally combine them into a complete contract strategy.
[0079] Similarly, the Critic network receives the state information of the current environment, including the supply and demand situation in the market, such as the market price fluctuation range of goods and the predicted value of market demand; and receives the contract strategy selected by the Actor network, that is, the action information. This includes the purchase price range determined in the contract, the delivery time window, the strictness of the quality acceptance standard, the specific requirements of after-sales service, etc. These action information will also be converted into appropriate numerical representations and input into the network. The input state and action information first enter the hidden layer of the Critic network. In the hidden layer, the network performs feature extraction operations on the input information. Each neuron will perform a weighted sum of part of the input information and process it through a non-linear activation function (such as ReLU) to extract more meaningful features. As the information is continuously transmitted and processed in the hidden layer, the features extracted by different neurons will gradually fuse. Through the calculation of multiple hidden layers, the network will synthesize various features to form a comprehensive understanding of the current state-action pair. Finally, the output layer of the network will output a single numerical value, which represents the expected long-term value (also known as value estimation) of adopting this contract strategy in the current state. For example, if the current market demand is strong and the contract strategy is reasonable price and short delivery time, the Critic network may output a higher value estimation, indicating that this strategy is expected to bring higher benefits in the long run. After the contract is executed, the platform will obtain an actual reward according to the actual revenue situation. According to the difference between the predicted value and the actual reward, the Critic network will use an optimization algorithm (such as gradient descent method) to update its own parameters. The goal is to make the predicted value of the network more accurately reflect the actual long-term revenue. For example, if the predicted value is always higher than the actual reward, the network will adjust the connection weights between neurons to reduce the future predicted value; otherwise, it will increase the predicted value. The evaluation result of the Critic network will be fed back to the Actor network. The Actor network will adjust the way of generating contract strategies according to the feedback of the Critic network. If the Critic network believes that the value of a certain strategy is high, the Actor network will increase the probability of adopting this strategy in a similar state; if the value is low, it will reduce the corresponding probability, so as to gradually learn the contract strategy that can obtain higher long-term revenue.
[0080] Exemplarily, a multi-agent reinforcement learning method based on the Actor-Critic architecture is adopted to improve the game strategies of insurance companies and reinsurers in contract negotiation: Actor network: Generate the optimal contract strategy (such as reinsurance ratio, reinsurance rate). Critic network: Evaluate the performance of the current strategy in long-term revenue and adjust the strategy.
[0081] S204: Initialize the neural network of each agent, the ReplayBuffer for storing the historical interaction data of the agent, and set the learning rate and discount factor.
[0082] S205: Each agent interacts with the environment based on the current policy, and obtains the state, selects an action, immediate reward, and the next state, and stores the interaction results in the Replay Buffer.
[0083] It should be understood that the interaction results include the state, action, reward, and the next state.
[0084] S206: Randomly sample from the Replay Buffer for training, calculate the Q-value update of each agent, and update the policy of the agent using the deterministic policy gradient, and use the target network for Q-value estimation; when the policy of the agent converges after multiple training cycles and the long-term return of each agent reaches the optimum, the training process terminates.
[0085] Specifically, the policy is updated according to the following formula:
[0086]
[0087] In the formula, is the learning rate, is the state-action value function, is with respect to the gradient of, is with respect to the gradient of.
[0088] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0089] In some embodiments, the dynamic adjustment and real-time feedback mechanism aims to ensure that the reinsurance contract can be adaptively optimized when the market, risk, and regulation change. This embodiment emphasizes the real-time adjustment and dynamic feedback of the contract parameters, and through the multi-agent cooperation mechanism, combines the impact of reinforcement learning on contract optimization to achieve a long-term stable optimal contract strategy.
[0090] 1. Reinforcement learning-driven feedback mechanism: During the contract adjustment process, this embodiment adopts a reinforcement learning-driven dynamic feedback mechanism to ensure that the contract terms can be optimized in real time under different market environments. The decision-making of the agent is not only based on the current market environment, but also optimized and adjusted according to the real-time feedback to ensure that the contract is always in the optimal state.
[0091] (1) Reinforcement learning fine-tuning: The core of reinforcement learning is to learn the optimal strategy. However, due to the constantly changing market environment, when the agent executes the game strategy, it needs to fine-tune the strategy to adapt to the new market situation. The contract terms after each adjustment will enter the feedback loop and optimize the decision-making in the following ways: Fast convergence: Reinforcement learning fine-tunes the strategy to avoid over-adjustment and ensure quickly reaching the optimal contract terms.
[0092] Exploration and exploitation balance: When the agent selects the adjustment strategy, according to the current market feedback, it chooses to explore (try new strategies) or exploit (optimize based on existing strategies).
[0093] (2) Real-time feedback on contract adjustment: After each contract adjustment, the agent will real-time correct the strategy according to the actual market feedback (such as changes in the insurance company's profit, changes in the reinsurer's claim risk, changes in the capital adequacy ratio, etc.). This process includes: Short-term feedback: Immediately monitor the market performance after the contract adjustment and adjust the reinsurance ratio, rate, or deductible.
[0094] Long-term feedback: In the long-term market environment, monitor the effect of the contract strategy and optimize the adjustment process through reinforcement learning.
[0095] 2. Reinforcement learning and agent collaboration: The reinforcement learning in this embodiment is not limited to the optimization of a single agent, but emphasizes the collaborative optimization of multiple agents. The actions of each agent will affect the decisions of other agents, thus achieving the overall optimal strategy of the system.
[0096] (1) Collaborative decision-making mechanism: The risk assessment agent calculates the future risk and provides adjustment suggestions to the insurance company and the reinsurer, affecting the contract parameters.
[0097] The regulatory compliance agent monitors the solvency capital requirement (SCR) and reminds the insurance company and the reinsurer to adjust the contract strategy to ensure compliance with regulatory requirements.
[0098] The external environment agent adjusts the market behavior according to factors such as macroeconomic changes and market risks, and provides market adaptation suggestions to other agents through the feedback mechanism.
[0099] (2) Real-time collaborative adjustment: Reinforcement learning optimizes the strategy of each agent through multi-agent collaboration. In each round of contract adjustment, the behaviors of each agent (such as the adjustment range and adjustment timing of the contract terms) will be collaboratively optimized to ensure the stability of the market and the long-term optimality of the contract.
[0100] 3. Stability and Compliance of Contract Adjustment: Stability during the contract optimization process is crucial, especially when the market fluctuates significantly. The dynamic adjustment mechanism ensures the stability and compliance of contract adjustment in the following ways: Smooth Transition: Contract parameter adjustments are gradual to avoid drastic fluctuations. Through the smooth update mechanism of reinforcement learning, it is ensured that the adjusted contract terms can transition smoothly and gradually adapt to the new market environment.
[0101] Compliance Assurance: Through the real-time monitoring of the regulatory compliance agent, it is ensured that the contract always complies with the latest regulatory requirements (such as capital adequacy ratio, solvency, etc.) during the adjustment process.
[0102] Compared with the existing technology, it has the following significant beneficial effects: 1. Achieve dynamic optimization of contract terms: By adopting the multi-agent reinforcement learning method, it is possible to adjust reinsurance contract terms (such as indemnity amount, deductible, risk sharing ratio, etc.) in real time according to dynamic information such as market changes, disaster risk assessment, and historical claim data, so as to ensure that reinsurers and insurance companies can always make optimal decisions in the changing risk environment. The self-adaptability of reinforcement learning enables the optimization process of contract terms to be automated without manual intervention, thus improving efficiency and reducing human bias.
[0103] 2. Multi-agent collaboration and game optimization: By designing a multi-agent system, different roles (such as reinsurers, insurance companies, disaster risk assessment agents, etc.) can collaborate and play games to achieve an optimal balance among multiple interests. This collaborative mechanism can effectively reduce the risks brought by the decision-making mistakes of a single agent and improve the overall efficiency of the system at the same time. Introducing a game mechanism among agents enables each agent to find the best strategy between cooperation and competition, effectively improving the stability and robustness of the system.
[0104] 3. Improve risk management capabilities: Through the real-time monitoring and assessment of factors such as disaster risks and market changes, the risk exposure can be dynamically adjusted, and the risk allocation in the reinsurance contract can be optimized. This enables reinsurers to better manage and diversify risks, thereby reducing the occurrence of potential losses. Different from the traditional static contract design, the present invention can adjust the contract terms in a timely manner before a disaster event occurs and continuously optimize based on real-time feedback, improving the flexibility and response speed of risk management.
[0105] 4. Improve decision-making efficiency and automation level: Through the self - learning and optimization of the reinforcement learning agent, it can quickly adjust decisions according to environmental changes, significantly improving the decision - making efficiency for optimizing reinsurance contract terms. The automation level of the system has been greatly improved, reducing the need for manual intervention, thus lowering labor costs and avoiding human decision - making errors.
[0106] 5. Adapt to the changing market environment: The multi - agent reinforcement learning system has strong adaptability and can cope with complex market environments and fluctuations in disaster risks. When market conditions change or disaster risk events occur, the system can promptly adjust decisions based on new information to ensure that the contract terms always adapt to environmental changes. It can make adaptive adjustments for different types of disaster events (such as storms, earthquakes, etc.) and optimize risk - sharing strategies, enhancing the accuracy and applicability of insurance products.
[0107] 6. Scalability and flexibility: It is possible to expand different agent roles according to actual needs to adapt to the optimization requirements of reinsurance contracts of different scales and types. For example, it is possible to add agent roles such as disaster risk assessment agents, different types of insurance companies, etc. as needed to ensure that the system can play its maximum effectiveness in multi - party collaboration. The framework is flexible and can be adjusted and optimized for different insurance products and reinsurance needs, applicable to a wide range of financial risk management fields.
[0108] 7. Enhanced transparency and interpretability: Through the combination of reinforcement learning and game theory, the decision - making process has high transparency, and the basis and motivation behind each step of the decision can be clearly traced. It can not only output the optimal contract terms but also provide the decision - making basis and risk assessment process, helping relevant decision - makers understand and accept the optimization results of the system.
[0109] As Figure 3 shown, the following is an embodiment of the optimization system for catastrophic reinsurance contracts provided by the present disclosure. It belongs to the same inventive concept as the optimization methods for catastrophic reinsurance contracts in the above - mentioned embodiments. For the details not described in detail in the embodiment of the optimization system for catastrophic reinsurance contracts, reference can be made to the embodiments of the optimization methods for catastrophic reinsurance contracts.
[0110] The optimization system for catastrophic reinsurance contracts includes: A construction unit for constructing multi - agents; A design unit for using a mathematical optimization model and combining the core terms in the reinsurance contract to design a dynamically optimized reinsurance contract mechanism to optimize reinsurance contract terms. Among them, optimizing the reinsurance contract mechanism includes reinsurance contract parameter optimization strategies and multi - agent collaborative optimization of contract parameters; Optimization unit, which is used to design the cooperation and game mechanisms among agents, and adopts the multi-agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real time. Among them, the cooperation and game mechanisms among agents include the game mechanism between the insurance company and the reinsurer, and the cooperation mechanism of agents.
[0111] Figure 4 It is a schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0112] The optimization method of the catastrophe reinsurance contract provided by the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0113] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0114] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0115] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0116] Among them, the processor may be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0117] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store the instructions or data just used or recycled by the processor. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0118] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0119] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0120] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modem processor, a baseband processor, etc.
[0121] The wireless communication module can provide solutions for wireless communications applied to electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0122] The electronic device can implement audio functions through an audio module, speaker, receiver, microphone, headphone jack, application processor, etc.
[0123] The electronic device can implement a shooting function through an ISP, camera, video codec, GPU, display screen, and application processor, etc.
[0124] The electronic device can implement a display function through a GPU, display screen, and application processor, etc.
[0125] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0126] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0127] In the storage medium provided in this application, there is a program product capable of implementing the optimization method of the catastrophe reinsurance contract.
[0128] The optimization method of the catastrophe reinsurance contract includes: constructing multiple agents; adopting a mathematical optimization model, combining the core terms in the reinsurance contract, and designing a dynamically optimized reinsurance contract mechanism to optimize the reinsurance contract terms. Among them, optimizing the reinsurance contract mechanism includes a reinsurance contract parameter optimization strategy and a multi-agent collaborative optimization of contract parameters; designing a cooperation and game mechanism among the agents, and adopting a multi-agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real time. Among them, the cooperation and game mechanism among the agents includes a game mechanism between the insurance company and the reinsurer, and a collaborative mechanism of the agents.
[0129] In some possible embodiments, the optimization method and system for a catastrophe reinsurance contract, which is the subject matter of the present disclosure, may be implemented in the form of a program product including program code that, when the program product runs on a terminal device, is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.
[0130] The storage medium of the present disclosure may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0131] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An optimization method for a catastrophe reinsurance contract, characterized in that, Including: Constructing multi - agents; Adopting a mathematical optimization model and combining the core terms in the reinsurance contract to design a dynamically optimized reinsurance contract mechanism to optimize the reinsurance contract terms. Among them, optimizing the reinsurance contract mechanism includes a reinsurance contract parameter optimization strategy and a multi - agent collaborative optimization of contract parameters; Designing the cooperation and game mechanism among agents, and using the multi - agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real - time. Among them, the cooperation and game mechanism among agents includes the game mechanism between the insurance company and the reinsurer, and the collaborative mechanism of agents.
2. The optimization method of the catastrophe reinsurance contract according to claim 1, wherein The constructed multi - agents at least include an insurance company agent, a reinsurer agent, a risk assessment agent, a regulatory compliance agent, and an external environment agent.
3. The optimization method of the catastrophe reinsurance contract according to claim 1, wherein The reinsurance contract parameter optimization strategy includes: cession ratio optimization, reinsurance rate optimization, and deductible optimization; Among them, the facultative reinsurance ratio is optimized according to the following formula ( ) In the formula, is the optimal reinsurance ratio, represents the penalty term when the capital adequacy ratio fails to meet the regulatory requirements, is the penalty coefficient, weighing the profit against the SCR violation risk, is the expected profit of the product line. Optimizing the reinsurance rate according to the following formula: In the formula, P* is the optimal reinsurance rate, E[L|L>d] is the expected loss after exceeding the deductible d, and ζ is the profit loading factor of the reinsurer. Optimizing the deductible according to the following formula: In the formula, is the optimal deductible, is the degree of risk exposure, is the risk adjustment parameter.
4. The optimization method of the catastrophe reinsurance contract according to claim 2, wherein The multi - agent collaborative optimization of contract parameters includes: For the insurance company agent, adjusting the cession ratio based on the capital adequacy ratio and market conditions; For the reinsurer agent, adjusting the reinsurance rate and deductible according to the market competition situation; For the regulatory compliance agent, providing real - time compliance checks to ensure that the capital adequacy ratio meets regulatory requirements; For the risk assessment agent, providing market risk forecasts to influence the decisions on the cession ratio and deductible; For the external environment agent, monitoring the market economic situation to influence the rate adjustment.
5. The optimization method of the catastrophe reinsurance contract according to claim 1, characterized in that The game mechanism between the insurance company and the reinsurer includes: Under stable market conditions, the insurance company and the reinsurer optimize the contract through Stackelberg reinforcement learning. The reinsurer first formulates the reinsurance contract parameters, and its goal is to maximize profits. Among them, the reinsurance contract parameters include the reinsurance rate, cession ratio, and deductible; After the reinsurer sets the contract parameters, the insurance company selects the optimal reinsurance purchase strategy based on its own capital adequacy ratio and market risk; When the market fluctuates violently or a disaster event occurs, the insurance company and the reinsurer choose to cooperate and share the benefits according to the Shapley value allocation mechanism.
6. The optimization method of the catastrophe reinsurance contract according to claim 1, wherein The multi - agent reinforcement learning method includes: When each agent makes a decision, defining a state vector for the market environment, contract terms, and historical data; For the insurance company agent, setting the cession ratio and whether to accept the reinsurance contract; for the reinsurer agent, setting the reinsurance rate and deductible; for the risk assessment agent, whether to renegotiate the reinsurance contract terms; for the regulatory compliance agent, whether to change the minimum solvency adequacy ratio; for the external environment agent, whether to change the risk exposure; Designing an agent reward function and using the Actor - Critic architecture to optimize the contract parameters. Among them, using the Actor network to generate the optimal contract strategy, using the Critic network to evaluate the performance of the current strategy in the long - term benefits, and adjusting the strategy; Initialize the neural network of each agent, the Replay Buffer for storing the historical interaction data of the agent, and set the learning rate and discount factor; Each agent interacts with the environment based on the current policy, obtains the state, selects an action, immediate reward, and the next state, and stores the interaction results in the Replay Buffer; Randomly sample from the Replay Buffer for training, calculate the Q-value update of each agent, update the agent's policy using the deterministic policy gradient, and use the target network for Q-value estimation; when the agent's policy converges after multiple training cycles and the long-term return of each agent reaches the optimum, the training process terminates.
7. The optimization method of the catastrophe reinsurance contract according to claim 6, characterized in that Perform policy update according to the following formula: Wherein, is the learning rate, is the state-action value function, is the gradient with respect to , is the gradient with respect to .
8. An optimization system for a catastrophe reinsurance contract, characterized in that, Including: A construction unit for constructing multi-agents; A design unit for using a mathematical optimization model and combining the core terms in the reinsurance contract to design a dynamically optimized reinsurance contract mechanism to optimize the reinsurance contract terms. Among them, optimizing the reinsurance contract mechanism includes a reinsurance contract parameter optimization strategy and a multi-agent collaborative optimization of contract parameters; An optimization unit for designing the cooperation and game mechanism between agents, and using the multi-agent reinforcement learning method to dynamically optimize the game strategy to adjust the contract terms in real time. Among them, the cooperation and game mechanism between agents includes the game mechanism between the insurance company and the reinsurer, and the cooperation mechanism of the agents.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the optimization method of the catastrophe reinsurance contract according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the optimization method of the catastrophe reinsurance contract according to any one of claims 1 to 7.
Citation Information
Patent Citations
Typhoon disaster assessment system
CN105741037A
Investment-reinsurance decision-making method and system
CN112925830A
Electricity market analogue simulation system, method and equipment and storage medium
CN117875153A
Reinsurance pricing calculation method based on multi-factor model
CN119273485A
Reinsurance and risk management method
US20020046066A1
Cited By
Reinforcement learning driven insurance underwriting evaluation method and system, and medium
CN121304355A