Intelligent premium pricing method and system based on multi-dimensional big data analysis
Through multi-dimensional big data analysis, the multi-head attention mechanism and long-term memory network optimized premium pricing method are used to solve the problem of insufficient data utilization in traditional methods and achieve the rationality and accuracy of premium pricing.
Patent Information
- Application Number
- CN202510958149.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional insurance pricing methods are difficult to make full use of massive and multi-dimensional data resources, resulting in unreasonable premium pricing and affecting the interests of insurance companies and buyers.
A cross-dimensional information cross-fusion method based on the multi-head attention mechanism is adopted, combining multiple sequences and time information, and data sequences are optimized through long-term and short-term memory networks, long-term dependencies are captured, key features are focused, and premium pricing prediction is carried out.
The rationality and accuracy of premium pricing are achieved, the ability to identify important influencing factors is improved, the long-term prediction ability is enhanced, and the accuracy of pricing recommendations is optimized.
Smart Images

Figure CN120450795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent price prediction, and specifically relates to an intelligent premium pricing method and system based on multi-dimensional big data analysis. Background Art
[0002] The insurance industry has traditionally used actuarial models for pricing, but with the improvement of data acquisition capabilities, traditional methods are unable to fully utilize massive, multi-dimensional data resources, which in turn leads to unreasonable premium pricing and affects the interests of insurance companies or insurance buyers. Summary of the Invention
[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an intelligent premium pricing method and system based on multi-dimensional big data analysis. In view of the problem that traditional methods are difficult to fully utilize massive, multi-dimensional data resources, which leads to unreasonable premium pricing, the present invention creatively adopts a cross-dimensional information cross-fusion method based on a multi-head attention mechanism. According to the big data historical data sequence of premium influencing factors and the corresponding premium pricing, a multivariate sequence is constructed. By performing cross-dimensional attention cross-fusion on the multivariate sequence, the mutual dependence between different influencing factors is processed through the channel attention mechanism, and on this basis, time information cross-fusion is performed to better capture the long-term dependence in multivariate data. By combining multivariate related information and time dependency information, the whole Comprehensively consider the multivariate influence and time correlation between premium prices and their influencing data, and then accurately predict and recommend premium pricing through influencing data to achieve reasonable premium pricing; the present invention incorporates patches, projections and global attention optimization based on long and short-term memory networks into the cross-dimensional information cross-fusion method based on the multi-head attention mechanism, performs patch division and projection on multivariate sequences, enhances the fitting ability of original data sequences, captures long-term dependencies of historical sequences of data through long and short-term memory networks, and retains important historical information, focuses on key features through the attention mechanism, improves the recognition ability of important influencing factors, and increases the prediction ability of long-term sequences on the basis of retaining global and local dependencies, thereby optimizing the accuracy of pricing prediction recommendations.
[0004] The technical solution provided by the present invention is as follows: a method for intelligent premium pricing based on multi-dimensional big data analysis, specifically comprising the following steps:
[0005] Step S1: Data collection, collecting historical series of premium pricing and historical series of data on related influencing factors;
[0006] Step S2: Data cleaning: clean all historical data sequences and remove outliers;
[0007] Step S3: Multivariate sequence construction, constructing the historical series of premium pricing and the historical series of its related influencing factor data into a multivariate sequence ,in , It is the sum of premium pricing and its related influencing factors;
[0008] Step S4: Normalize the sequence using variance and mean values for the multivariate sequence Normalize and get the normalized sequence : ;
[0009] Where, represents the variance, represents the average value, is a preset very small constant, is the normalized sequence;
[0010] Step S5: Patch construction and projection, divide the normalized sequence into multiple patches, and expand the dimension of each patch through the projection layer to obtain the expanded sequence : ;
[0011] Where, represents the patch partitioning operation, represents the projection operation, which is implemented by a multi-layer perceptron, That is, the extended sequence, whose dimension is P×N, P represents the dimension after projection, and N represents the total number of patches;
[0012] Step S6: Global attention optimization based on long short-term memory network to expand the sequence Perform global attention optimization based on long-term and short-term networks to obtain the optimized sequence , specifically including the following steps:
[0013] Step S61: Long short-term memory network processing, using a long short-term memory network with N hidden layers to process the expanded sequence Process and get N hidden layer outputs ;
[0014] Step S62: Calculate the attention score function. Calculate the attention score function corresponding to the time step through the hidden layer output: ;
[0015] Where, and represents the weight coefficient, represents the bias coefficient, represents the hyperbolic tangent activation function, Represents the hidden layer output The corresponding attention score function;
[0016] Step S63: Calculate the attention score according to the attention score function: ;
[0017] Where, represents the attention score, represents the exponential function, is the preset adjustment parameter;
[0018] Step S64: Global attention mechanism output, based on all calculated attention scores and their corresponding hidden layer outputs Perform weighted summation to obtain the global attention output;
[0019] Step S65: Activate the output and process the global attention output with activation function to obtain the optimized sequence ;
[0020] Step S7: Based on the cross-dimensional information cross-fusion of the multi-head attention mechanism, the optimized sequence is cross-fused based on the multi-head attention mechanism to obtain sequence features, which specifically includes the following steps:
[0021] Step S71: Add position embedding, add learnable position embedding to the optimized sequence along the time dimension, and obtain the sequence : ;
[0022] Where, represents a learnable position embedding that preserves the temporal order of the original optimization sequence;
[0023] Step S72: Channel attention calculation, for each step size n in each attention head in the multi-head attention, calculate the query, key and value matrix 、 and : ; ; ;
[0024] Where, 、 and represents the learnable weight matrix, 、 and represents the query, key, and value matrices, Representative An attention head, Representative sequence Middle The matrix transpose of the step sequence, n ∈ [1, N ] , i.e. each patch represents a time step;
[0025] Step S73: Calculate the channel attention output to obtain the channel attention output: ;
[0026] Where, represents the softmax activation function, Representative Channel attention output of a step sequence;
[0027] Step S74: Multi-head output splicing, splicing the channel attention outputs of all attention heads and linearly projecting them to obtain the channel representation ;
[0028] Step S75: Time attention calculation based on channel representation, for each channel representation in the multi-head attention Perform temporal attention processing to obtain query, key, and value matrices 、 and : ; ; ;
[0029] Where, 、 and represents the learnable weight matrix, 、 and represents the query, key, and value matrices, Representative channel representation Middle The matrix transpose of the sequence corresponding to the channels, m ∈ [1, M ] , that is, each influencing factor data represents a channel;
[0030] Step S76: Temporal attention output calculation, and concatenation and linear projection of the temporal attention outputs of all attention heads are performed to obtain sequence features. The temporal attention output calculation formula is as follows: ;
[0031] Step S8: Prediction output, input the sequence features into the feedforward layer and activate the output to obtain a multivariate prediction sequence. Finally, based on the multivariate prediction sequence and the current premium pricing-related influencing factor data, the current premium prediction pricing, i.e., the recommended pricing, is obtained.
[0032] The premium intelligent pricing system based on multi-dimensional big data analysis provided by the present invention includes a data acquisition module, a data cleaning module, a sequence processing module and an intelligent prediction pricing module;
[0033] The data collection module collects historical series of premium pricing and historical series of data on related influencing factors;
[0034] The data cleaning module cleans all historical data sequences and removes outliers.
[0035] The sequence processing module constructs the historical sequence of premium pricing and the historical sequence of its related influencing factor data into a multivariate sequence, and performs sequence normalization and dimension expansion to obtain an expanded sequence;
[0036] The intelligent prediction pricing module adopts global attention optimization based on long short-term memory network to perform global attention optimization based on long short-term network on the extended sequence, and adopts cross-dimensional information cross-fusion based on multi-head attention mechanism to perform information fusion and prediction output to obtain recommended pricing.
[0037] The beneficial results achieved by the present invention using the above scheme are as follows:
[0038] (1) The present invention creatively adopts a cross-dimensional information cross-fusion method based on a multi-head attention mechanism. According to the big data historical data sequence of premium influencing factors and the corresponding premium pricing, a multivariate sequence is constructed. By performing cross-dimensional attention cross-fusion on the multivariate sequence, the mutual dependence between different influencing factors is processed through the channel attention mechanism, and on this basis, time information cross-fusion is performed to better capture the long-term dependence in the multivariate data. By combining multivariate related information and time-dependent information, the multivariate influence and time correlation between premium prices and their influencing data are comprehensively considered, and then the premium pricing is accurately predicted and recommended through the influencing data, thereby achieving reasonable premium pricing;
[0039] (2) The present invention integrates patching, projection and global attention optimization based on long short-term memory network into the cross-dimensional information cross-fusion method based on multi-head attention mechanism, performs patch division and projection on multivariate sequences, enhances the fitting ability of original data sequences, captures the long-term dependencies of historical sequences of data through long short-term memory network, and retains important historical information. It focuses on key features through attention mechanism, improves the recognition ability of important influencing factors, and increases the prediction ability of long-term sequences on the basis of retaining global and local dependencies, thereby optimizing the accuracy of pricing prediction recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A flow chart of the intelligent premium pricing method based on multi-dimensional big data analysis provided by the present invention;
[0041] Figure 2 This is a module diagram of the premium intelligent pricing system based on multi-dimensional big data analysis provided by the present invention.
[0042] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0044] Example 1, see Figure 1 The present invention provides an intelligent premium pricing method based on multi-dimensional big data analysis, which specifically includes the following steps:
[0045] Step S1: Data collection, collecting historical series of premium pricing and historical series of data on related influencing factors;
[0046] Step S2: Data cleaning: clean all historical data sequences and remove outliers;
[0047] Step S3: Multivariate sequence construction, constructing the historical series of premium pricing and the historical series of its related influencing factor data into a multivariate sequence ,in , It is the sum of premium pricing and its related influencing factors;
[0048] Step S4: Normalize the sequence using variance and mean values for the multivariate sequence Normalize and get the normalized sequence : ;
[0049] Where, represents the variance, represents the average value, is a preset very small constant, is the normalized sequence;
[0050] Step S5: Patch construction and projection, divide the normalized sequence into multiple patches, and expand the dimension of each patch through the projection layer to obtain the expanded sequence : ;
[0051] Where, represents the patch partitioning operation, represents the projection operation, which is implemented by a multi-layer perceptron, That is, the extended sequence, whose dimension is P×N, P represents the dimension after projection, and N represents the total number of patches;
[0052] Step S6: Global attention optimization based on long short-term memory network to expand the sequence Perform global attention optimization based on long-term and short-term networks to obtain the optimized sequence ;
[0053] Step S7: Based on the cross-dimensional information cross-fusion of the multi-head attention mechanism, the optimized sequence is cross-fused based on the multi-head attention mechanism to obtain sequence features;
[0054] Step S8: Prediction output, input the sequence features into the feedforward layer and activate the output to obtain a multivariate prediction sequence. Finally, based on the multivariate prediction sequence and the current premium pricing-related influencing factor data, the current premium prediction pricing, i.e., the recommended pricing, is obtained.
[0055] Embodiment 2: This embodiment is based on the above embodiment, and step S6 specifically includes the following steps:
[0056] Step S61: Long short-term memory network processing, using a long short-term memory network with N hidden layers to process the expanded sequence Process and get N hidden layer outputs ;
[0057] Step S62: Calculate the attention score function. Calculate the attention score function corresponding to the time step through the hidden layer output: ;
[0058] Where, and represents the weight coefficient, represents the bias coefficient, represents the hyperbolic tangent activation function, Represents the hidden layer output The corresponding attention score function;
[0059] Step S63: Calculate the attention score according to the attention score function: ;
[0060] Where, represents the attention score, represents the exponential function, is the preset adjustment parameter;
[0061] Step S64: Global attention mechanism output, based on all calculated attention scores and their corresponding hidden layer outputs Perform weighted summation to obtain the global attention output;
[0062] Step S65: Activate the output and process the global attention output with activation function to obtain the optimized sequence .
[0063] Embodiment 3: This embodiment is based on the above embodiment, and step S7 specifically includes the following steps:
[0064] Step S71: Add position embedding, add learnable position embedding to the optimized sequence along the time dimension, and obtain the sequence : ;
[0065] Where, represents a learnable position embedding that preserves the temporal order of the original optimization sequence;
[0066] Step S72: Channel attention calculation, for each step n in each attention head in the multi-head attention, calculate the query, key and value matrix 、 and : ; ; ;
[0067] Where, 、 and represents the learnable weight matrix, 、 and represents the query, key, and value matrices, Representative An attention head, Representative sequence Middle The matrix transpose of the step sequence, n ∈ [1, N ] , i.e. each patch represents a time step;
[0068] Step S73: Calculate the channel attention output to obtain the channel attention output: ;
[0069] Where, represents the softmax activation function, Representative Channel attention output of a step sequence;
[0070] Step S74: Multi-head output splicing, splicing the channel attention outputs of all attention heads and linearly projecting them to obtain the channel representation ;
[0071] Step S75: Temporal attention calculation based on channel representation, for each channel representation in the multi-head attention Perform temporal attention processing to obtain query, key, and value matrices 、 and : ; ; ;
[0072] Where, 、 and represents the learnable weight matrix, 、 and represents the query, key, and value matrices, Representative channel representation Middle The matrix transpose of the sequence corresponding to the channels, m ∈ [1, M ] , that is, each influencing factor data represents a channel;
[0073] Step S76: Temporal attention output calculation, and concatenation and linear projection of the temporal attention outputs of all attention heads are performed to obtain sequence features. The temporal attention output calculation formula is as follows: .
[0074] Example 4, see Figure 2 This embodiment is based on the above embodiment. The premium intelligent pricing system based on multi-dimensional big data analysis provided by the present invention includes a data acquisition module, a data cleaning module, a sequence processing module and an intelligent prediction pricing module;
[0075] The data collection module collects historical series of premium pricing and historical series of data on related influencing factors;
[0076] The data cleaning module cleans all historical data sequences and removes outliers.
[0077] The sequence processing module constructs the historical sequence of premium pricing and the historical sequence of its related influencing factor data into a multivariate sequence, and performs sequence normalization and dimension expansion to obtain an expanded sequence;
[0078] The intelligent prediction pricing module adopts global attention optimization based on long short-term memory network to perform global attention optimization based on long short-term network on the extended sequence, and adopts cross-dimensional information cross-fusion based on multi-head attention mechanism to perform information fusion and prediction output to obtain recommended pricing.
[0079] Example 5: This example is based on the above example, and the technology of the present invention is implemented as follows:
[0080] (1) Data collection and preprocessing
[0081] Data Source
[0082] Internal data: historical underwriting data, claims data, and customer behavior data;
[0083] External data: third-party credit data, geographic location data, social media data, meteorological data, etc.
[0084] (2) Data preprocessing
[0085] Data cleaning: remove outliers and noise;
[0086] Feature engineering: extract relevant features and the relationship between features;
[0087] Data integration: Use primary keys to associate and integrate multi-source heterogeneous data;
[0088] (3) Multi-dimensional analysis
[0089] Utilizes a global attention optimization method based on long short-term memory networks to capture long-term historical dependencies of multivariate sequences;
[0090] Utilize cross-dimensional information fusion based on a multi-head attention mechanism to perform information fusion and prediction output to obtain recommended pricing;
[0091] (4) Smart pricing strategy
[0092] Develop personalized premium pricing strategies based on multi-dimensional data analysis results:
[0093] Pricing influencing factors:
[0094] Customer behavior factors: driving behavior, health status;
[0095] Environmental factors: accident-prone areas, weather impacts;
[0096] Historical factors: historical claims frequency, default record;
[0097] Dynamic adjustment: Real-time monitoring of key factors and dynamic adjustment of premiums.
[0098] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0099] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0100] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. An intelligent premium pricing method based on multi-dimensional big data analysis, characterized by: The specific steps include: Step S1: Collect historical series of premium pricing and historical series of data on related influencing factors; Step S2: Clean the historical series of all data to remove outliers; Step S3: Construct the historical series of premium pricing and its related influencing factor data into a multivariate series ,in , It is the sum of premium pricing and its related influencing factors; Step S4: Use variance and mean value to analyze the multivariate series Normalize and get the normalized sequence ; Step S5: Divide the normalized sequence into multiple patches, and expand the dimension of each patch through the projection layer to obtain the expanded sequence ; Step S6: Extend the sequence Perform global attention optimization based on long-term and short-term networks to obtain the optimized sequence ; Step S7: Perform cross-dimensional information cross-fusion on the optimized sequence based on the multi-head attention mechanism to obtain sequence features; Step S8: Input the sequence features into the feedforward layer and activate the output to obtain a multivariate prediction sequence. Finally, based on the multivariate prediction sequence and the current premium pricing-related influencing factor data, the current premium prediction pricing, i.e., the recommended pricing, is obtained.
2. The method for intelligent premium pricing based on multi-dimensional big data analysis according to claim 1, characterized in that: Step S6 specifically includes the following steps: Step S61: Use a long short-term memory network with N hidden layers to expand the sequence Process and get N hidden layer outputs ; Step S62: Calculate the attention score function corresponding to the time step through the hidden layer output: ; Where, and represents the weight coefficient, represents the bias coefficient, represents the hyperbolic tangent activation function, Represents the hidden layer output The corresponding attention score function; Step S63: Calculating the attention score according to the attention score function; Step S64: Calculate the attention scores and their corresponding hidden layer outputs based on all Perform weighted summation to obtain the global attention output; Step S65 performs activation function processing on the global attention output to obtain the optimized sequence .
3. The method for intelligent premium pricing based on multi-dimensional big data analysis according to claim 2, characterized in that: Step S7 specifically includes the following steps: Step S71: Add learnable position embedding to the optimized sequence along the time dimension to obtain the sequence ; Step S72: Calculate the query, key, and value matrices for each step size n in each attention head in the multi-head attention 、 and ; Step S73: Calculate the channel attention output to obtain the channel attention output: ; Where, represents the softmax activation function, Representative Channel attention output of a step sequence; Step S74: Concatenate and linearly project the channel attention outputs of all attention heads to obtain the channel representation ; Step S75: For the channel representation of each attention head in the multi-head attention Perform temporal attention processing to obtain query, key, and value matrices 、 and ; Step S76: Calculate the temporal attention output, and concatenate and linearly project the temporal attention outputs of all attention heads to obtain sequence features.
4. An intelligent premium pricing system based on multi-dimensional big data analysis, used to implement the intelligent premium pricing method based on multi-dimensional big data analysis as described in any one of claims 1 to 3, characterized in that: It includes data acquisition module, data cleaning module, sequence processing module and intelligent prediction pricing module; The data collection module collects historical series of premium pricing and historical series of data on related influencing factors; The data cleaning module cleans all historical data sequences and removes outliers. The sequence processing module constructs the historical sequence of premium pricing and the historical sequence of its related influencing factor data into a multivariate sequence, and performs sequence normalization and dimension expansion to obtain an expanded sequence; The intelligent prediction pricing module adopts global attention optimization based on long short-term memory network to perform global attention optimization based on long short-term network on the extended sequence, and adopts cross-dimensional information cross-fusion based on multi-head attention mechanism to perform information fusion and prediction output to obtain recommended pricing.