Insurance policy generation method, system, terminal and medium based on reinforcement learning
Through the insurance strategy generation method based on reinforcement learning, the data acquisition strategy is dynamically adjusted using multimodal data and deep learning models, and the problem of difficult to eliminate inverse selection risks in the existing technology is solved, and more scientific and adaptable insurance services are achieved.
Patent Information
- Application Number
- CN202510134026.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
In the face of complex market environment and diversified customer behavior patterns, the existing technology cannot effectively eliminate the risk of inverse selection, resulting in an increase in the overall risk level of insurance companies underwriting customers.
The insurance strategy generation method based on reinforcement learning is adopted, and by establishing authorized connections between users and insurance companies, multi-modal data is obtained, and deep learning models and reinforcement learning models are built, the data acquisition frequency and priority are dynamically adjusted to generate accurate insurance strategies.
It has achieved accurate data acquisition, optimized risk assessment and insurance strategy allocation, improved the scientificity and adaptability of insurance services, and reduced the impact of reverse selection risks.
Smart Images

Figure CN119558992B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of insurance risk assessment, and specifically relates to a method, system, terminal and medium for generating an insurance strategy based on reinforcement learning. Background Art
[0002] With the rapid development of the insurance industry, customer adverse selection risk has gradually become one of the major challenges facing the industry. Adverse selection risk refers to the fact that due to information asymmetry, high-risk customers are more inclined to buy insurance products, while low-risk customers may choose to participate less or completely exit the market.
[0003] Adverse selection risk directly leads to an increase in the overall risk level of insurance companies' insured customers, which in turn affects pricing strategies, profitability, and the effectiveness of risk management. Current risk assessment methods mainly rely on historical information provided by customers to classify risks, and then match corresponding insurance policies based on risk assessment results.
[0004] Although these methods provide technical support for risk control to a certain extent, due to the subjectivity of customers, they are unable to eliminate the impact of adverse selection risks when faced with complex market environments and diverse customer behavior patterns. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides an insurance policy generation method, system, terminal and medium based on reinforcement learning to solve the above-mentioned technical problems.
[0006] In a first aspect, the present invention provides a method for generating an insurance policy based on reinforcement learning, comprising:
[0007] S1, establishing an authorized connection relationship between the user and the insurance company, the insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority, the multimodal data including the user's health data, consumption behavior data and social interaction data;
[0008] S2, building a deep learning model to train a risk assessment model based on multimodal data within multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data;
[0009] S3, configures insurance policies for users based on user status and risk probability, and trains a reinforcement learning model based on user status, risk probability and insurance policies. The input of the model is user status and risk probability, and the output of the model is insurance policies;
[0010] S4, obtaining user status changes and user status change frequency based on the predicted status, adjusting data collection frequency based on the status change combined with the status threshold and risk probability, adjusting collection priority based on the risk probability of each type of data, and adjusting the preset time window based on the user status change frequency;
[0011] S5, within the adjusted time window, the current multimodal data is obtained based on the adjusted data collection frequency and collection priority, the current multimodal data is input into the risk assessment model to obtain the current predicted user status and risk probability, and the current predicted user status and risk probability are input into the reinforcement learning model to obtain the current insurance policy.
[0012] In an optional implementation, step 2 specifically includes:
[0013] Pre-build the objective function based on user status and risk probability:
[0014]
[0015] in, and is the weight parameter for balancing the prediction error of the behavior state and the prediction error of the risk probability, is the control regularization term The parameter of influence, N is the total number of data types, is the user status of the i-th data type, is the predicted user status of the i-th data type, is the risk probability of the jth data type, is the predicted risk probability of the jth data type;
[0016] After preprocessing, the multimodal data is divided into training set, validation set and test set;
[0017] Initialize the parameters of the risk assessment model, input the multimodal data in the training set into the model, and the model reaches the output layer after calculation of the hidden layer, outputting the predicted status and risk probability of each type of data;
[0018] The loss value of the model is calculated based on the objective function. The gradient of the objective function with respect to each parameter of the model is calculated based on the loss value combined with the chain rule, and the back propagation algorithm is used to update the parameters.
[0019] Verify the model based on the validation set and calculate the model's performance indicators. When the model's performance indicators are stable, the training is completed.
[0020] Use the test set to evaluate the trained model to obtain the evaluation results, and adjust the model parameters based on the evaluation results to obtain the final risk assessment model.
[0021] In an optional implementation, in step S4, adjusting the collection priority based on the risk probability of each type of data specifically includes:
[0022]
[0023] in, is the priority weight of the i-th type of data, is the predicted risk probability of the i-th type of data, and N is the total number of data types;
[0024] Sort the priority weights of all types of data from large to small to obtain the adjusted collection priority order.
[0025] In an optional embodiment, when adjusting the collection priority, the predicted risk probability of each data type is judged against its corresponding high risk threshold, and the types corresponding to the predicted risk probabilities exceeding the high risk threshold are collected in real time and are not included in the priority sorting.
[0026] In an optional implementation, in step S4, the acquisition frequency adjustment specifically includes:
[0027] Calculate the state-change adjusted acquisition frequency based on the current state change of each type of data:
[0028]
[0029] in, is the weight of the i-th data, is the basic acquisition frequency of the ith type of data, is the acquisition frequency after the state change of the i-th data is adjusted, is the adjustment coefficient for the i-th data type, is the user status change of the i-th data type, is the threshold parameter of the i-th data type;
[0030] Calculate the adjusted acquisition priority based on the predicted risk probability for each type of data, and calculate the predicted risk-adjusted acquisition frequency based on the adjusted acquisition priority:
[0031] ;
[0032] The collection frequency adjusted for weighted comprehensive status change and the collection frequency adjusted for predicted risk are then used to obtain the comprehensive adjusted collection frequency.
[0033] In an optional implementation, in step S3, the training of the reinforcement learning model specifically includes:
[0034] Get the profit value of the insurance company, and design the reward function based on the profit value at each moment combined with the risk probability:
[0035]
[0036] in, is the profit value of the insurance company. It is the risk penalty factor, which is used to weigh the relationship between return and risk. is the overall risk level;
[0037] S3-1, taking the insurance company as an intelligent agent, defines the state space based on user status and risk probability, and defines the action space based on insurance policy;
[0038] S3-2, obtain the current user status and risk probability, and select an insurance strategy from the action space based on the current user status and risk probability combined with the greedy strategy;
[0039] S3-3, after the agent executes the selected insurance strategy, it calculates the corresponding reward based on the reward function, and updates the pros and cons of the corresponding insurance strategy based on the current user status, risk probability, reward, and the next user status and risk probability;
[0040] S3-4, repeat the processes of S3-2 and S3-3 to make the insurance strategy selected by the agent converge to the optimal one.
[0041] In an optional implementation, step S3-3 specifically includes:
[0042] Based on the Q-learning algorithm, the strategy value function between the environment state and the insurance strategy is constructed. The Q value is updated based on the current user state, risk probability, reward, and the next user state and risk probability to update the pros and cons between the current environment state and the insurance strategy. Specifically:
[0043]
[0044] in, yes, is the environmental state composed of the user state and risk probability at time t, is the insurance strategy at time t, is the learning rate, is the immediate reward at time t, is the discount factor, It is the optimal future strategy value obtained by taking the optimal insurance strategy under the environmental state at time t+1. It's an insurance strategy. It is the optimal insurance strategy adopted under the environmental conditions at time t+1.
[0045] In a second aspect, the present invention provides an insurance policy generation system based on reinforcement learning. When the system is implemented, the above-mentioned insurance policy generation method based on reinforcement learning is executed. The system includes:
[0046] The data collection module establishes an authorized connection relationship between the user and the insurance company. The insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority. The multimodal data includes the user's health data, consumption behavior data, and social interaction data;
[0047] The risk assessment model building module builds a deep learning model and trains the risk assessment model based on multimodal data in multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data;
[0048] The reinforcement learning model building module configures insurance policies for users based on user status and risk probability, and trains the reinforcement learning model based on user status, risk probability and insurance policy. The input of the model is user status and risk probability, and the output of the model is insurance policy.
[0049] The data collection adjustment module obtains the user status change and the frequency of the user status change based on the predicted status, adjusts the data collection frequency based on the status change combined with the status threshold and the risk probability, adjusts the collection priority based on the risk probability of each type of data, and adjusts the preset time window based on the frequency of the user status change;
[0050] The current policy acquisition module acquires the current multimodal data within the adjusted time window based on the adjusted data collection frequency and collection priority, inputs the current multimodal data into the risk assessment model to obtain the current predicted user status and risk probability, and inputs the current predicted user status and risk probability into the reinforcement learning model to obtain the current insurance policy.
[0051] In a third aspect, a terminal is provided, including:
[0052] processor, memory, wherein:
[0053] The memory is used to store computer programs.
[0054] The processor is used to call and run the computer program from the memory, so that the terminal executes the above-mentioned terminal method.
[0055] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on a computer, the computer executes the methods described in the above aspects.
[0056] The beneficial effects of the present invention lie in that the insurance policy generation method, system, terminal and medium based on reinforcement learning provided by the present invention acquire multimodal data by establishing an authorized connection between users and insurance companies, build and train risk assessment and reinforcement learning models to configure insurance policies, and dynamically adjust data collection frequency, priority and time window according to predicted status, risk probability, etc., so as to achieve accurate data acquisition, optimize risk assessment and insurance policy configuration, improve the scientificity and adaptability of insurance services, and reduce the impact of adverse selection risks.
[0057] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the description will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0059] Figure 1 It is a schematic flow chart of a method for generating an insurance policy based on reinforcement learning according to an embodiment of the present invention.
[0060] Figure 2 It is a schematic block diagram of an insurance policy generation system based on reinforcement learning according to an embodiment of the present invention.
[0061] Figure 3 A schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0064] The insurance policy generation method based on reinforcement learning provided in the embodiment of the present invention is executed by a computer device, and accordingly, the insurance policy generation system based on reinforcement learning runs in the computer device.
[0065] Figure 1 is a schematic flow chart of a method for generating an insurance policy based on reinforcement learning according to an embodiment of the present invention. Figure 1 The execution subject may be an insurance policy generation system based on reinforcement learning. According to different requirements, the order of the steps in the flowchart may be changed, and some may be omitted.
[0066] like Figure 1 As shown, the method includes:
[0067] S1, establishing an authorized connection relationship between the user and the insurance company, the insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority, the multimodal data including the user's health data, consumption behavior data and social interaction data;
[0068] An authorized connection channel is established between users and insurance companies, so that both parties can establish a trust relationship and exchange data. Insurance companies collect multimodal data covering health data, consumer behavior data and social interaction data from users within a specific preset time window based on the pre-set basic collection frequency and collection priority. These data reflect the user's living status and behavior patterns from different dimensions, providing a basis for subsequent analysis.
[0069] S2, building a deep learning model to train a risk assessment model based on multimodal data within multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data;
[0070] The risk assessment model is built using deep learning related technologies, and the multimodal data collected in multiple preset time windows are input into the model as training samples. The model deeply analyzes and learns these data to explore the potential patterns and rules behind the data, so that it can predict the future state of each type of data and give the corresponding risk probability. Through the comprehensive analysis of multimodal data, taking into account multiple aspects of user life, it can predict risks more comprehensively and accurately compared with the evaluation method of a single data type.
[0071] S3, configures insurance policies for users based on user status and risk probability, and trains a reinforcement learning model based on user status, risk probability and insurance policies. The input of the model is user status and risk probability, and the output of the model is insurance policies;
[0072] Based on the user status and risk probability output by the risk assessment model, combined with the insurance business expertise and market demand, an insurance strategy is tailored for the user. Afterwards, the user status and risk probability are used as input, and the insurance strategy is used as output to train the reinforcement learning model so that the model can learn the optimal insurance strategy configuration method under different risk conditions.
[0073] S4, obtaining user status changes and user status change frequency based on the predicted status, adjusting data collection frequency based on the status change combined with the status threshold and risk probability, adjusting collection priority based on the risk probability of each type of data, and adjusting the preset time window based on the user status change frequency;
[0074] Monitor the changes in user status and the frequency of changes based on the status predicted by the risk assessment model. Compare the status changes with the pre-set status thresholds, and dynamically adjust the data collection frequency based on the risk probability. Adjust the collection priority accordingly based on the risk probability of each type of data. According to the frequency of user status changes, reasonably adjust the preset time window to adapt to the rhythm of data changes.
[0075] S5, within the adjusted time window, the current multimodal data is obtained based on the adjusted data collection frequency and collection priority, the current multimodal data is input into the risk assessment model to obtain the current predicted user status and risk probability, and the current predicted user status and risk probability are input into the reinforcement learning model to obtain the current insurance policy.
[0076] Within the adjusted time window, the user's current multimodal data is obtained again according to the optimized data collection frequency and collection priority. The latest data is input into the trained risk assessment model to obtain the current predicted user status and risk probability. Then, these latest prediction results are input into the reinforcement learning model to obtain the latest insurance strategy that adapts to the current user status.
[0077] Optionally, as an embodiment of the present invention, based on providing users with insurance discounts, etc., users are allowed to authorize their health data (including physical examination reports, medical history, exercise volume, etc.), consumption behavior data (purchase categories, shopping frequency, payment methods), and social interaction data (likes, comments, shares on social media, etc.) to insurance companies. The data obtained is only used to assess risks and generate insurance policies, and does not involve the user's specific privacy information.
[0078] Optionally, as an embodiment of the present invention, a dedicated feature extraction network is designed for data types of different modalities to obtain key features of each modality. For example, for health data, a convolutional neural network (CNN) is used to extract key trends in time series; for behavioral data, a multi-head attention mechanism is used to capture patterns in transaction frequency and amount distribution; for social data, a Transformer model is used to extract sentiment tendencies and behavioral features in text.
[0079] By introducing a modality-aware network, the importance score of each modality in the current task is dynamically calculated, and the weight of each modality is automatically adjusted. Then, the multi-head attention mechanism is used to calculate the interactive features between modalities and capture the implicit correlation between different modalities. Finally, the features of each modality are weighted and fused to generate the final fusion features.
[0080] Optionally, as an embodiment of the present invention, step 2 specifically includes:
[0081] Pre-build the objective function based on user status and risk probability:
[0082]
[0083] in, and is the weight parameter for balancing the prediction error of the behavior state and the prediction error of the risk probability, is the control regularization term The parameter of influence, N is the total number of data types, is the user status of the i-th data type, is the predicted user status of the i-th data type, is the risk probability of the jth data type, is the predicted risk probability of the jth data type;
[0084]
[0085] Among them, v is the regularization coefficient and W is the weight parameter of the model.
[0086] The fusion features obtained after preprocessing the multimodal data are used as the data set, and the data set is divided into a training set, a validation set, and a test set;
[0087] Initialize the parameters of the risk assessment model, input the multimodal data in the training set into the model, and the model reaches the output layer after calculation of the hidden layer, outputting the predicted status and risk probability of each type of data;
[0088] The loss value of the model is calculated based on the objective function. The gradient of the objective function with respect to each parameter of the model is calculated based on the loss value combined with the chain rule, and the back propagation algorithm is used to update the parameters.
[0089] Verify the model based on the validation set and calculate the model's performance indicators. When the model's performance indicators are stable, the training is completed.
[0090] Use the test set to evaluate the trained model to obtain the evaluation results, and adjust the model parameters based on the evaluation results to obtain the final risk assessment model.
[0091] Optionally, as an embodiment of the present invention, the model outputs the risk probability of the customer:
[0092]
[0093] in, and are model parameters, The risk probability for the customer.
[0094] Based on risk probability, users are classified into "high risk", "medium risk" and "low risk".
[0095] Optionally, as an embodiment of the present invention, in step S3, the training of the reinforcement learning model specifically includes:
[0096] Get the profit value of the insurance company, and design the reward function based on the profit value at each moment combined with the risk probability:
[0097]
[0098] in, is the profit value of the insurance company. It is the risk penalty factor, which is used to weigh the relationship between return and risk. is the overall risk level;
[0099] S3-1, taking the insurance company as an intelligent agent, defines the state space based on the user state and risk probability, and defines the action space based on the insurance policy; the agent's strategy determines which action to choose in each state. Usually, at the beginning, the strategy is random, that is, the agent chooses any action in the action space with the same probability in each state.
[0100] S3-2, obtain the current user status and risk probability, and select an insurance strategy from the action space based on the current user status and risk probability combined with the greedy strategy;
[0101] S3-3, after the agent executes the selected insurance strategy, it calculates the corresponding reward based on the reward function, and updates the pros and cons of the corresponding insurance strategy based on the current user status, risk probability, reward, and the next user status and risk probability;
[0102] S3-4, repeat the processes of S3-2 and S3-3 to make the insurance strategy selected by the agent converge to the optimal one.
[0103] Optionally, as an embodiment of the present invention, step S3-3 specifically includes:
[0104] Based on the Q-learning algorithm, the strategy value function between the environment state and the insurance strategy is constructed. The Q value is updated based on the current user state, risk probability, reward, and the next user state and risk probability to update the pros and cons between the current environment state and the insurance strategy. Specifically:
[0105]
[0106] in, yes, is the environmental state composed of the user state and risk probability at time t, is the insurance strategy at time t, is the learning rate, is the immediate reward at time t, is the discount factor, It is the optimal future strategy value obtained by taking the optimal insurance strategy under the environmental state at time t+1. It's an insurance strategy. It is the optimal insurance strategy adopted under the environmental conditions at time t+1.
[0107] Q-learning is an algorithm based on reinforcement learning that is used to find the optimal strategy in a given environment. It helps the agent make the best decision by learning a Q function (action value function) to evaluate the long-term cumulative reward of taking a specific action in a specific state.
[0108] Optionally, as an embodiment of the present invention, there is an example in which the insurance company includes the driving behavior data (such as emergency braking frequency, speeding times, daily mileage) of the car owner A, vehicle information (vehicle age, vehicle model, vehicle value), historical claims records (number of claims, claim amount) and traffic conditions in the area (accident rate, congestion level) into the state space. Specifically, A has a 3-year-old family car, emergency braking frequency 5 times a month, speeding times 3 times a year, mileage 15,000 kilometers per year, 1 small claim record, medium accident rate and severe congestion in the area, and these information together constitute the state of A at a certain moment.
[0109] The action space includes various insurance policy adjustment options, including adjusting premiums (which can be set to increase or decrease the basic premium by 10%, 20%, 30%, etc.), adjusting the amount of insurance (for example, increasing or decreasing the amount of insurance within a certain range), and changing insurance terms (such as increasing or decreasing certain specific protection items, such as whether to include glass breakage insurance, spontaneous combustion insurance, etc.).
[0110] Insurance companies design reward functions based on their own profit goals and risk assessments. If A has good driving behavior within an insurance period, the accident risk is reduced, the insurance company's profits increase, and the reward value is positive; conversely, if A has frequent accidents and the claims expenses are large, the reward value is negative.
[0111] The Q value table is randomly initialized, the learning rate is set to 0.1, and the discount factor is set to 0.9. At this time, the Q value is random for the state of A and various insurance strategy action combinations.
[0112] Based on the greedy strategy, in the initial stage, the insurance company may randomly choose an action, such as keeping the current premium and insured amount unchanged and only adding glass breakage insurance.
[0113] After executing the selected strategy, the reward is calculated based on A's actual situation during the insurance period. If A does not have an accident during the period and the insurance company's profit increases, assuming that the profit value increases by 500 yuan, the overall risk level is evaluated to be reduced, and the reward value calculated according to the reward function is 500-0.5*100=450 (assuming that the risk penalty factor is 0.5 and the risk level is quantified as 100).
[0114] Use the Q-learning algorithm formula to update the Q value. Assuming that the maximum Q value when taking the optimal action under the next environmental state is 500, and the initial Q value of the current environmental state and insurance strategy is 100, then the updated Q value = 100 + 0.1 * [450 + 0.9 * 500-100] = 180.
[0115] By repeating the above process, as the number of training times increases, the insurance company can gradually learn the optimal strategy for different states. After a large number of training cycles, such as 1,000 trainings, the model gradually converges and obtains the optimal insurance strategy for car owners such as A. This may be to reduce the basic premium by 10%, increase the insurance amount by a certain amount, and reasonably adjust the protection items according to the situation in the region and the characteristics of the vehicle.
[0116] Optionally, as an embodiment of the present invention, in step S4, adjusting the collection priority based on the risk probability of each type of data specifically includes:
[0117]
[0118] in, is the priority weight of the i-th type of data, is the predicted risk probability of the i-th type of data, and N is the total number of data types;
[0119] Sort the priority weights of all types of data from large to small to obtain the adjusted collection priority order.
[0120] Optionally, as an embodiment of the present invention, when adjusting the collection priority, the predicted risk probability of each data type is judged respectively with its corresponding high risk threshold, and the types corresponding to the predicted risk probabilities exceeding the high risk threshold are collected in real time and are not included in the priority sorting.
[0121] The predicted risk probability of each data type is judged against its corresponding medium risk threshold. For the types corresponding to the predicted risk probability exceeding the medium risk threshold but not exceeding the high risk threshold, their collection priority is increased:
[0122]
[0123] in, is the priority weight after triggering, To trigger the weight adjustment coefficient.
[0124] Optionally, as an embodiment of the present invention, adjusting the preset time window based on the user status change frequency specifically includes:
[0125]
[0126] in, is the maximum time window, is the adjustment factor.
[0127] In some embodiments, the insurance policy generation system based on reinforcement learning may include multiple functional modules composed of computer program segments. The computer program of each program segment in the insurance policy generation system based on reinforcement learning may be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Functionality for reinforcement learning-based insurance policy generation.
[0128] In this embodiment, the insurance policy generation system based on reinforcement learning can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system may include: data acquisition module, risk assessment model building module, reinforcement learning model building module, data acquisition adjustment module, current strategy acquisition module. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes:
[0129] The data collection module establishes an authorized connection relationship between the user and the insurance company. The insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority. The multimodal data includes the user's health data, consumption behavior data, and social interaction data;
[0130] The risk assessment model building module builds a deep learning model and trains the risk assessment model based on multimodal data in multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data;
[0131] The reinforcement learning model building module configures insurance policies for users based on user status and risk probability, and trains the reinforcement learning model based on user status, risk probability and insurance policy. The input of the model is user status and risk probability, and the output of the model is insurance policy.
[0132] The data collection adjustment module obtains the user status change and the frequency of the user status change based on the predicted status, adjusts the data collection frequency based on the status change combined with the status threshold and the risk probability, adjusts the collection priority based on the risk probability of each type of data, and adjusts the preset time window based on the frequency of the user status change;
[0133] The current policy acquisition module acquires the current multimodal data within the adjusted time window based on the adjusted data collection frequency and collection priority, inputs the current multimodal data into the risk assessment model to obtain the current predicted user status and risk probability, and inputs the current predicted user status and risk probability into the reinforcement learning model to obtain the current insurance policy.
[0134] An authorized connection is established through the data collection module to obtain user multimodal data. The risk assessment model construction module uses multimodal data training to obtain the predicted status and risk probability. The reinforcement learning model construction module configures the insurance policy and trains the model accordingly. The data collection adjustment module flexibly adjusts the collection frequency, priority and time window according to the predicted status. The current policy acquisition module obtains the current data and updates the insurance policy. The overall efficient cycle of data collection and analysis, strategy formulation and adjustment is realized, which effectively improves the accuracy of insurance companies' risk assessment, the scientific nature of strategy formulation, and the dynamic adaptability to changes in user status, and enhances the risk management capabilities and operational efficiency of insurance business.
[0135] Figure 3 A schematic diagram of the structure of a terminal provided in an embodiment of the present invention, wherein the terminal can be used to execute the method for generating insurance policies based on reinforcement learning provided in an embodiment of the present invention.
[0136] The terminal may include: a processor, a memory and a communication unit. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. It may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0137] The memory can be used to store the execution instructions of the processor, and the memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory are executed by the processor, the terminal is enabled to execute some or all of the steps in the following method embodiments.
[0138] The processor is the control center of the storage terminal, which uses various interfaces and lines to connect various parts of the entire electronic terminal, and executes various functions and / or processes data of the electronic terminal by running or executing software programs and / or modules stored in the memory, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor can only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.
[0139] The communication unit is used to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals or send user data to other terminals.
[0140] The present invention also provides a computer-readable storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0141] Therefore, the technical effects that can be achieved by this embodiment can be found in the description above and will not be repeated here.
[0142] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, including several instructions for enabling a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0143] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the terminal embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.
[0144] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, which can be electrical, mechanical or other forms.
[0145] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0147] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.
Claims
1. A method for generating insurance policies based on reinforcement learning, characterized in that: The following steps are involved: S1, establishing an authorized connection relationship between the user and the insurance company, the insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority, the multimodal data including the user's health data, consumption behavior data and social interaction data; S2, building a deep learning model to train a risk assessment model based on multimodal data within multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data; Step 2 specifically includes: Pre-build the objective function based on user status and risk probability: in, and is the weight parameter for balancing the prediction error of the behavior state and the prediction error of the risk probability, is the control regularization term The parameter of influence, N is the total number of data types, is the user status of the i-th data type, is the predicted user status of the i-th data type, is the risk probability of the jth data type, is the predicted risk probability of the jth data type; After preprocessing, the multimodal data is divided into training set, validation set and test set; Initialize the parameters of the risk assessment model, input the multimodal data in the training set into the model, and the model reaches the output layer after calculation of the hidden layer, outputting the predicted status and risk probability of each type of data; The loss value of the model is calculated based on the objective function. The gradient of the objective function with respect to each parameter of the model is calculated based on the loss value combined with the chain rule, and the back propagation algorithm is used to update the parameters. Verify the model based on the validation set and calculate the model's performance indicators. When the model's performance indicators are stable, the training is completed. Use the test set to evaluate the trained model to obtain the evaluation results, adjust the model parameters based on the evaluation results, and obtain the final risk assessment model; S3, configures insurance policies for users based on user status and risk probability, and trains a reinforcement learning model based on user status, risk probability and insurance policies. The input of the model is user status and risk probability, and the output of the model is insurance policies; S4, obtaining user status changes and user status change frequency based on the predicted status, adjusting data collection frequency based on the status change combined with the status threshold and risk probability, adjusting collection priority based on the risk probability of each type of data, and adjusting the preset time window based on the user status change frequency; S5, within the adjusted time window, the current multimodal data is obtained based on the adjusted data collection frequency and collection priority, the current multimodal data is input into the risk assessment model to obtain the current predicted user status and risk probability, and the current predicted user status and risk probability are input into the reinforcement learning model to obtain the current insurance policy.
2. The insurance policy generation method based on reinforcement learning according to claim 1, characterized in that: In step S4, adjusting the collection priority based on the risk probability of each type of data specifically includes: in, is the priority weight of the i-th type of data, is the predicted risk probability of the i-th type of data, and N is the total number of data types; Sort the priority weights of all types of data from large to small to obtain the adjusted collection priority order.
3. The insurance policy generation method based on reinforcement learning according to claim 2, characterized in that: When adjusting the collection priority, the predicted risk probability of each data type will be judged against its corresponding high-risk threshold. The types corresponding to the predicted risk probabilities exceeding the high-risk threshold will be collected in real time and will not be included in the priority sorting.
4. The insurance policy generation method based on reinforcement learning according to claim 1, characterized in that: In step S4, the acquisition frequency adjustment specifically includes: Calculate the state-change adjusted acquisition frequency based on the current state change of each type of data: in, is the weight of the i-th data, is the basic acquisition frequency of the ith type of data, is the acquisition frequency after the state change of the i-th data is adjusted, is the adjustment coefficient for the i-th data type, is the user status change of the i-th data type, is the threshold parameter of the i-th data type; Calculate the adjusted acquisition priority based on the predicted risk probability for each type of data, and calculate the predicted risk-adjusted acquisition frequency based on the adjusted acquisition priority: ; The collection frequency adjusted for weighted comprehensive status changes and the collection frequency adjusted for predicted risk are then used to obtain the comprehensive adjusted collection frequency.
5. The insurance policy generation method based on reinforcement learning according to claim 1, characterized in that: In step S3, the training of the reinforcement learning model specifically includes: Get the profit value of the insurance company, and design the reward function based on the profit value at each moment combined with the risk probability: in, is the profit value of the insurance company. It is the risk penalty factor, which is used to weigh the relationship between return and risk. is the overall risk level; S3-1, taking the insurance company as an intelligent agent, defines the state space based on user status and risk probability, and defines the action space based on insurance policy; S3-2, obtain the current user status and risk probability, and select an insurance strategy from the action space based on the current user status and risk probability combined with the greedy strategy; S3-3, after the agent executes the selected insurance strategy, it calculates the corresponding reward based on the reward function, and updates the pros and cons of the corresponding insurance strategy based on the current user status, risk probability, reward, and the next user status and risk probability; S3-4, repeat the processes of S3-2 and S3-3 to make the insurance strategy selected by the agent converge to the optimal one.
6. The insurance policy generation method based on reinforcement learning according to claim 5, characterized in that: Step S3-3 specifically includes: Based on the Q-learning algorithm, the strategy value function between the environment state and the insurance strategy is constructed. The Q value is updated based on the current user state, risk probability, reward, and the next user state and risk probability to update the pros and cons between the current environment state and the insurance strategy. Specifically: in, yes, is the environmental state composed of the user state and risk probability at time t, is the insurance strategy at time t, is the learning rate, is the immediate reward at time t, is the discount factor, It is the optimal future strategy value obtained by taking the optimal insurance strategy under the environmental state at time t+1. It's an insurance strategy. It is the optimal insurance strategy adopted under the environmental conditions at time t+1.
7. An insurance policy generation system based on reinforcement learning, characterized in that: When the system is implemented, the method for generating an insurance policy based on reinforcement learning as described in any one of claims 1 to 6 is executed, and the system includes: The data collection module establishes an authorized connection relationship between the user and the insurance company. The insurance company obtains the user's multimodal data within a preset time window based on the basic collection frequency and collection priority. The multimodal data includes the user's health data, consumption behavior data, and social interaction data; The risk assessment model building module builds a deep learning model and trains the risk assessment model based on multimodal data in multiple preset time windows. The input of the risk assessment model is multimodal data, and the output is the predicted status and risk probability of each type of data; The reinforcement learning model building module configures insurance policies for users based on user status and risk probability, and trains the reinforcement learning model based on user status, risk probability and insurance policy. The input of the model is user status and risk probability, and the output of the model is insurance policy. The data collection adjustment module obtains the user status change and the frequency of the user status change based on the predicted status, adjusts the data collection frequency based on the status change combined with the status threshold and the risk probability, adjusts the collection priority based on the risk probability of each type of data, and adjusts the preset time window based on the frequency of the user status change; The current policy acquisition module acquires the current multimodal data within the adjusted time window based on the adjusted data collection frequency and collection priority, inputs the current multimodal data into the risk assessment model to obtain the current predicted user status and risk probability, and inputs the current predicted user status and risk probability into the reinforcement learning model to obtain the current insurance policy.
8. A terminal, characterized in that: include: A memory, used for storing an insurance policy generation program based on reinforcement learning; A processor, configured to implement the steps of the insurance policy generation method based on reinforcement learning as described in any one of claims 1 to 6 when executing the insurance policy generation program based on reinforcement learning.
9. A computer-readable storage medium, characterized in that: The readable storage medium stores an insurance policy generation program based on reinforcement learning, and when the insurance policy generation program based on reinforcement learning is executed by a processor, the steps of the insurance policy generation method based on reinforcement learning as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Vehicle insurance risk prediction method, vehicle insurance risk prediction device, and server
CN107292528A
Risk assessment method and device
CN109670724A
Physical examination-free limit prediction method and device, electronic equipment and storage medium
CN114781687A