Insurance pricing method based on web crawler technology

Through network crawling technology collection and deep learning algorithms to analyze customer emotional data, integrate traditional insurance pricing models, solve the problem that traditional pricing methods fail to consider customer emotional and fraud identification lag, achieve more accurate risk assessment and personalized pricing, and improve customer satisfaction and fraud risk management.

CN119991188APending Publication Date: 2025-05-13SHANGHAI YUYANG EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510058819.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional insurance pricing methods fail to fully consider customers' subjective emotional state and risk preferences, resulting in inaccurate pricing, difficult to meet customers' personalized needs, and there is a problem of fraud identification lag.

Method used

The insurance pricing method based on web crawling technology is adopted, and customer emotional data is collected from social platforms, online insurance comment sections and customer service interaction records, combined with deep learning algorithms and professional emotional thesaurus, and sentiment analysis is carried out to transform emotional tendencies into risk preferences, and integrated them into traditional pricing models to build a dynamic pricing model.

Benefits of technology

It realizes more accurate customer risk assessment and personalized pricing solutions, improves customer satisfaction and loyalty, and identify potential signs of fraud in advance through real-time emotional monitoring to reduce the risk of insurance fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991188A_ABST
    Figure CN119991188A_ABST
Patent Text Reader

Abstract

The invention discloses an insurance pricing method based on a web crawler technology, and the method comprises the steps: automatically capturing customer emotion data from a plurality of social platforms, an online insurance comment region and a customer service interaction record through the web crawler technology, and carrying out the cleaning and preprocessing of the collected customer emotion data, obtaining an emotion data set for subsequent emotion analysis; a deep learning algorithm is adopted, a professional emotion word bank and weight setting are combined, accurate recognition and quantitative analysis of customer emotion expression are achieved, and the emotion tendency is accurately converted into the risk preference degree; factors such as customer risk preference obtained through sentiment analysis and the like are integrated into a traditional insurance pricing model, the weight of each factor in pricing is determined by applying mathematical modeling and statistical analysis methods, a dynamic pricing model is constructed, and insurance policy terms and insurance premium of customers are calculated and adjusted in real time; according to the method, more accurate insurance policy pricing is realized, product recommendation better meeting customer requirements is realized, and meanwhile, the risk prevention and control capability and the market competitiveness of a company are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to the insurance industry, and in particular to an insurance pricing method based on web crawler technology. Background Art

[0002] In the insurance industry, traditional policy pricing mainly relies on basic information about customers, such as age, gender, occupation, health status, etc. However, with the development of society and changes in consumer behavior, this pricing method has gradually shown its limitations. On the one hand, customers' emotional state and risk preferences have an important impact on their insurance needs and behaviors, but traditional pricing models fail to fully consider these factors. For example, a customer who is optimistic and positive about life may have different risk prevention awareness from a customer with a negative attitude, and their insurance needs and ability to bear premiums may also be different.

[0003] On the other hand, with the increasingly fierce market competition, customers' demand for personalized and customized insurance products is constantly increasing. Traditional pricing models are difficult to meet customers' diverse expectations, resulting in the impact on customer satisfaction and loyalty. At the same time, insurance fraud has also brought great challenges to the industry. Traditional risk assessment methods have certain lags and limitations in identifying potential fraudulent behaviors. Traditional pricing methods only consider objective risk factors, ignoring the impact of customers' subjective emotions and psychological factors on insurance demand and risk behavior. For example, two customers with the same objective risk characteristics (such as age and health status) may have different levels of demand for insurance and risk prevention behaviors due to different emotional states (one is optimistic and positive, the other is anxious and worried), but traditional pricing cannot distinguish them. Customers' emotional states may change over time, while traditional pricing models are relatively static and cannot reflect the impact of these changes on risks in a timely manner, resulting in pricing that does not match actual risks. Traditional risk assessment mainly focuses on customers' objective information and historical claims data, and the identification of fraudulent behaviors is mostly conducted post-review at the claims stage. However, fraudsters may circumvent traditional risk assessment by concealing real information before or during the insurance process. The lack of sentiment analysis technology makes it difficult to detect potential fraud signs in advance from customers' emotional expressions and behavioral patterns. Summary of the invention

[0004] In order to solve the defects of the prior art, the present invention provides an insurance pricing method based on network crawler technology.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] The present invention discloses an insurance pricing method based on web crawler technology, comprising the following steps: step 1, emotional data collection, using web crawler technology to automatically capture customer emotional data from multiple social platforms, online insurance review areas, and customer service interaction records, cleaning and preprocessing the collected customer emotional data, and obtaining an emotional data set for subsequent emotional analysis;

[0007] Step 2: Conduct sentiment analysis, using deep learning algorithms, combined with professional sentiment vocabulary and weight settings, to accurately identify and quantify customer sentiment expressions, and accurately convert sentiment tendencies into risk preference levels;

[0008] Step 3: Incorporate factors such as customer risk preferences obtained through sentiment analysis into the traditional insurance pricing model, use mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, build a dynamic pricing model, and calculate and adjust the customer's policy terms and premiums in real time.

[0009] As a preferred technical solution of the present invention, the method of cleaning and preprocessing the collected customer sentiment data in step 1 to obtain a sentiment data set for subsequent sentiment analysis is: assuming that the original collected text data set is D = {d 1 , d 2 , …, d n}, where d i Represents the i-th text data. These text data are segmented during the data collection phase to build a high-quality data set. A large amount of insurance field text data is used to train the topic model to obtain the topic-word distribution. And the document-topic distribution φ, for the text to be segmented d i , in the traditional word segmentation results Based on the information of the topic model, we first find two adjacent words w ij and w i(j+1) , calculate the probability P(w ij , w i(j+1) |z), where z represents the topic; if the probability is greater than the empirical coefficient 0.73, the two words are merged into one phrase.

[0010] As a preferred technical solution of the present invention, the step 2 adopts a deep learning algorithm, combined with a professional sentiment word library and weight setting, to achieve accurate recognition and quantitative analysis of customer sentiment expression, and the method of accurately converting sentiment tendency into risk preference degree is to design a deep learning network integrating attention mechanism, multimodal information and reinforcement learning feedback; the deep learning network is used for sentiment analysis and risk quantification judgment;

[0011] The overall structure of the deep learning network is divided into an input layer, a feature extraction layer, an attention mechanism layer, a multimodal fusion layer, a risk quantification layer, and a reinforcement learning feedback layer;

[0012] The input layer is used to receive text data, image data and voice data;

[0013] Among them, text data: the pre-processed insurance-related text data is converted into word vector representation, and the word vector is obtained using Word2Vec word embedding technology to form a sequence input;

[0014] Image data: The collected images such as insurance promotional posters and claims-related pictures are preprocessed to make them suitable for the input of convolutional neural networks.

[0015] Voice data: The collected voice signals such as customer service call recordings and customer feedback audio are converted into a processable form of spectrogram through methods such as short-time Fourier transform.

[0016] As a preferred technical solution of the present invention, the feature extraction layer is used to extract features from text data, image data and voice data:

[0017] The feature extraction layer extracts text features of text data by using a bidirectional long short-term memory network (Bi-LSTM) to process the text word vector sequence, which can simultaneously capture the forward and backward information of the text, effectively handle the long sequence dependency problem, and learn to obtain rich text semantic features;

[0018] The feature extraction layer extracts features of the image data by using a convolutional neural network ResNet50 to extract features of the image data, and learns common image features on a large-scale image dataset to capture key information related to insurance in the image;

[0019] The feature extraction layer extracts the speech features of the speech data by using a convolutional neural network to perform a convolution operation on the spectrum graph to extract the acoustic features of the speech.

[0020] As a preferred technical solution of the present invention, the attention mechanism layer applies the attention mechanism to the feature representation of text, image and speech respectively, and highlights the text part that is more important for sentiment analysis and risk quantification by calculating the attention weight between the hidden states of each time step; similarly, the image features and speech features are also processed by the attention mechanism to focus on key information; the detailed process is as follows:

[0021] Three different features are used: text features text features Where T is the length of the text sequence, d Tis the text feature dimension; image feature Where N is the number of image features, d I is the image feature dimension; speech feature Where M is the number of speech features, d S is the speech feature dimension; the above-mentioned features are obtained after being processed by the feature extraction layer of each modality. During feature transformation, we perform linear transformation on the features of each modality to make them of the same dimension for subsequent calculation; for text features T, we use the linear transformation matrix W T Until T′=TW T ,in d is the unit after unified measurement; for image feature I, through the linear transformation matrix W I We get I′=IW I ,in For the speech feature S, through the linear transformation matrix W S , we get S′SW S ,in By calculating the attention scores of these three features and normalizing each attention score matrix using the Softmax function, we can get the attention weight matrix:

[0022] Text-image attention weight matrix W TI :

[0023] Text-speech attention weight matrix W TS :

[0024] Image-speech attention weight matrix W IS : After the above cross-modal attention mechanism is processed, the text feature T that integrates different modal information is obtained. fused 、Image Features I fused and speech feature S fused The fused features will be used as the input of the subsequent multimodal fusion layer, and further integrated to generate a unified feature representation containing rich cross-modal correlation information.

[0025] As a preferred technical solution of the present invention, the risk quantification layer is used to perform nonlinear transformation on the basis of fusion features using a multi-layer perceptron to map the fusion features to the prediction space of sentiment scores and risk preferences; the sentiment scores are converted into probability distributions of different sentiment categories (positive, negative, neutral) through the Softmax layer, and the quantitative values ​​of risk preferences are output at the same time.

[0026] The specific process of integrating factors such as customer risk preferences obtained through sentiment analysis into the traditional insurance pricing model, using mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, building a dynamic pricing model, and calculating and adjusting the customer's policy terms and premiums in real time is as follows:

[0027] The risk preference level (R) of the customer is obtained through sentiment analysis. The value range is between ([0, 1]), (0) indicates a very high risk aversion level, and (1) indicates a very high risk preference level. In order to reflect the impact of sentiment analysis factors on pricing, a dynamic adjustment coefficient (k) is introduced;

[0028] (k=αR+β)

[0029] Among them, (α) and (β) are parameter dynamic pricing (P dynamic ) is: (P dynamic =P traditional ×(1+k))

[0030] Taking into account dynamic factors such as market competition and customer feedback, a market dynamic factor (M) is introduced, which is determined by factors such as market supply and demand, competitor pricing, etc., and its value range is between ([0.8, 1.2]). The final dynamic pricing formula is:

[0031] (P dynamic =P traditional ×(1+k)×M)

[0032] For product recommendations, we built a personalized insurance recommendation system based on customers' emotional data and behavioral data (combined with emotional factors) to recommend insurance products that best meet customers' needs and expectations.

[0033] As a preferred technical solution of the present invention, for the product recommendation part, we build a personalized insurance recommendation system based on the customer's emotional data and behavioral data, combined with emotional factors, to recommend the insurance products that best meet the customer's needs and expectations. The specific process is: first, obtain the emotional tendency vector (E) of customer (i) through emotional analysis i ) (e.g., numerical representations of sentiment dimensions such as positive, negative, and neutral), each insurance product \(j\) also has a corresponding sentiment adaptation vector (F j ) (determined according to product characteristics and target customers’ emotional needs);

[0034] Then calculate the matching degree between customer sentiment and product sentiment (match(i, j)):

[0035]

[0036] Next, the sentiment matching is integrated into the prediction score to obtain the final recommendation score.

[0037]

[0038] Among them, (ω∈[0,1]) is a weight coefficient used to balance the impact of the prediction score based on collaborative filtering and the sentiment matching degree. ij ) Sort products and recommend products with higher ratings to customers.

[0039] The beneficial effects of the present invention are:

[0040] This insurance pricing method based on web crawler technology introduces sentiment analysis and fully considers the impact of customer subjective emotional factors on insurance demand and risk behavior. Compared with the traditional pricing method that only relies on objective factors, it can more accurately assess customer risks, formulate more reasonable premium prices, and avoid customer loss or excessive company risks caused by unreasonable pricing; based on real-time collection and analysis of emotional data, combined with customer behavior data, it can accurately grasp the dynamic changes in customer needs and provide customers with personalized insurance product recommendations and customized pricing plans. Compared with traditional market research and customer behavior analysis methods, it is more timely and accurate, and effectively improves customer satisfaction and loyalty; with the help of sentiment analysis technology, it continuously monitors customer emotions and behavioral changes in all aspects of insurance business (before insurance, during insurance, and during insurance), and can detect potential fraud signs in advance, change the situation of lagging traditional fraud risk identification, reduce insurance fraud risks, and protect company interests; provide personalized and dynamically adjusted insurance products and services to meet the diverse needs of customers, so that the company can stand out in the fiercely competitive insurance market, attract more customers, expand market share, and promote sustainable development of the company. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0042] In the attached picture:

[0043] Figure 1 It is a flow chart of an insurance pricing method based on network crawler technology of the present invention. DETAILED DESCRIPTION

[0044] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0045] Example: Figure 1As shown, the present invention provides an insurance pricing method based on network crawler technology, comprising the following steps:

[0046] Step 1: Collect sentiment data: Use web crawler technology to automatically capture customer sentiment data from multiple social platforms, online insurance review areas, and customer service interaction records, clean and pre-process the collected customer sentiment data, and obtain a sentiment data set for subsequent sentiment analysis;

[0047] Step 2: Conduct sentiment analysis, using deep learning algorithms, combined with professional sentiment vocabulary and weight settings, to accurately identify and quantify customer sentiment expressions, and accurately convert sentiment tendencies into risk preference levels;

[0048] Step 3: Incorporate factors such as customer risk preferences obtained through sentiment analysis into the traditional insurance pricing model, use mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, build a dynamic pricing model, and calculate and adjust the customer's policy terms and premiums in real time. By collecting customer sentiment data through multiple channels, using advanced machine learning algorithms for precise analysis, and building a dynamic pricing model, more accurate policy pricing and product recommendations that better meet customer needs can be achieved, while also improving the company's risk prevention and control capabilities and market competitiveness.

[0049] Wherein, the method of cleaning and preprocessing the collected customer sentiment data in step 1 to obtain the sentiment data set for subsequent sentiment analysis is: assuming that the original collected text data set is D = {d 1 , d 2 , …, d n}, where d i Represents the i-th text data. These text data are segmented during the data collection phase to build a high-quality data set. A large amount of insurance field text data is used to train the topic model to obtain the topic-word distribution. And the document-topic distribution φ, for the text to be segmented d i , in the traditional word segmentation results Based on the information of the topic model, we first find two adjacent words w ij and w i(j+1) , calculate the probability P(w ij , w i(j+1) |z), where z represents the topic; if the probability is greater than the empirical coefficient 0.73, the two words are merged into one phrase.

[0050] In the step 2, a deep learning algorithm is used, and a professional sentiment word library and weight setting are combined to realize accurate recognition and quantitative analysis of customer sentiment expression. The method of accurately converting sentiment tendency into risk preference degree is to design a deep learning network integrating attention mechanism, multimodal information and reinforcement learning feedback; the deep learning network is used for sentiment analysis and risk quantification judgment;

[0051] The overall structure of the deep learning network is divided into an input layer, a feature extraction layer, an attention mechanism layer, a multimodal fusion layer, a risk quantification layer, and a reinforcement learning feedback layer;

[0052] The input layer is used to receive text data, image data and voice data;

[0053] Among them, text data: the pre-processed insurance-related text data is converted into word vector representation, and the word vector is obtained using Word2Vec word embedding technology to form a sequence input;

[0054] Image data: The collected images such as insurance promotional posters and claims-related pictures are preprocessed to make them suitable for the input of convolutional neural networks.

[0055] Voice data: The collected voice signals such as customer service call recordings and customer feedback audio are converted into a processable form of spectrogram through methods such as short-time Fourier transform.

[0056] The feature extraction layer is used to extract features from text data, image data, and voice data:

[0057] The feature extraction layer extracts text features of text data by using a bidirectional long short-term memory network (Bi-LSTM) to process the text word vector sequence, which can simultaneously capture the forward and backward information of the text, effectively handle the long sequence dependency problem, and learn to obtain rich text semantic features;

[0058] The feature extraction layer extracts features of the image data by using a convolutional neural network ResNet50 to extract features of the image data, and learns common image features on a large-scale image dataset to capture key information related to insurance in the image;

[0059] The feature extraction layer extracts the speech features of the speech data by using a convolutional neural network to perform a convolution operation on the spectrum graph to extract the acoustic features of the speech.

[0060] The attention mechanism layer applies attention mechanisms to the feature representations of text, images, and speech respectively, and highlights the text parts that are more important for sentiment analysis and risk quantification by calculating the attention weights between the hidden states at each time step. Similarly, the image features and speech features are also processed by the attention mechanism to focus on key information. The detailed process is as follows:

[0061] Three different features are used: text features text features Where T is the length of the text sequence, d T is the text feature dimension; image feature Where N is the number of image features, d I is the image feature dimension; speech feature Where M is the number of speech features, d S is the speech feature dimension; the above-mentioned features are obtained after being processed by the feature extraction layer of each modality. During feature transformation, we perform linear transformation on the features of each modality to make them of the same dimension for subsequent calculation; for text features T, we use the linear transformation matrix W T Until T′=TW T ,in d is the unit after unified measurement; for image feature I, through the linear transformation matrix W I We get I′=IW I ,in For the speech feature S, through the linear transformation matrix W S , we get S′SW S ,in By calculating the attention scores of these three features and normalizing each attention score matrix using the Softmax function, we can get the attention weight matrix:

[0062] Text-image attention weight matrix W TI :

[0063] Text-speech attention weight matrix W TS :

[0064] Image-speech attention weight matrix W IS : After the above cross-modal attention mechanism is processed, the text feature T that integrates different modal information is obtained. fused 、Image Features I fused and speech feature S fused The fused features will be used as the input of the subsequent multimodal fusion layer, and further integrated to generate a unified feature representation containing rich cross-modal correlation information.

[0065] The risk quantification layer is used to perform nonlinear transformation using a multi-layer perceptron based on the fusion features, and map the fusion features to the prediction space of sentiment scores and risk preferences; the sentiment scores are converted into probability distributions of different sentiment categories (positive, negative, neutral) through the Softmax layer, and the quantitative value of risk preferences is output at the same time. The specific process of integrating factors such as customer risk preferences obtained from sentiment analysis into the traditional insurance pricing model, using mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, constructing a dynamic pricing model, and calculating and adjusting the customer's policy terms and premiums in real time is as follows:

[0066] The risk preference level (R) of the customer is obtained through sentiment analysis. The value range is between ([0, 1]), (0) indicates a very high risk aversion level, and (1) indicates a very high risk preference level. In order to reflect the impact of sentiment analysis factors on pricing, a dynamic adjustment coefficient (k) is introduced;

[0067] (k=αR+β)

[0068] Among them, (α) and (β) are parameter dynamic pricing (P dynamic ) is: (P dynamic =P traditional ×(1+k))

[0069] Taking into account dynamic factors such as market competition and customer feedback, a market dynamic factor (M) is introduced, which is determined by factors such as market supply and demand, competitor pricing, etc., and its value range is between ([0.8, 1.2]). The final dynamic pricing formula is:

[0070] (P dynamic =P traditional ×(1+k)×M)

[0071] For the product recommendation part, we build a personalized insurance recommendation system based on the customer's emotional data and behavioral data (, combined with emotional factors, to recommend the insurance products that best meet their needs and expectations to customers. For the product recommendation part, we build a personalized insurance recommendation system based on the customer's emotional data and behavioral data (, combined with emotional factors, to recommend the insurance products that best meet their needs and expectations to customers. The specific process is: first, obtain the emotional tendency vector (E of customer (i)) through sentiment analysis i ) (e.g., numerical representations of sentiment dimensions such as positive, negative, and neutral), each insurance product \(j\) also has a corresponding sentiment adaptation vector (F j ) (determined according to product characteristics and target customers’ emotional needs);

[0072] Then calculate the matching degree between customer sentiment and product sentiment (match(i, j)):

[0073]

[0074] Next, the sentiment matching degree is integrated into the prediction score to obtain the final recommendation score (S ij ):

[0075]

[0076] Among them, (ω∈[0,1]) is a weight coefficient used to balance the impact of the prediction score based on collaborative filtering and the sentiment matching degree. ij ) Sort products and recommend products with higher ratings to customers.

[0077] This insurance pricing method based on web crawler technology introduces sentiment analysis and fully considers the impact of customer subjective emotional factors on insurance demand and risk behavior. Compared with the traditional pricing method that only relies on objective factors, it can more accurately assess customer risks, formulate more reasonable premium prices, and avoid customer loss or excessive company risks caused by unreasonable pricing; based on real-time collection and analysis of emotional data, combined with customer behavior data, it can accurately grasp the dynamic changes in customer needs and provide customers with personalized insurance product recommendations and customized pricing plans. Compared with traditional market research and customer behavior analysis methods, it is more timely and accurate, and effectively improves customer satisfaction and loyalty; with the help of sentiment analysis technology, it continuously monitors customer emotions and behavioral changes in all aspects of insurance business (before insurance, during insurance, and during insurance), and can detect potential fraud signs in advance, change the situation of lagging traditional fraud risk identification, reduce insurance fraud risks, and protect company interests; provide personalized and dynamically adjusted insurance products and services to meet the diverse needs of customers, so that the company can stand out in the fiercely competitive insurance market, attract more customers, expand market share, and promote sustainable development of the company.

[0078] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An insurance pricing method based on web crawler technology, characterized in that: The following steps are included: Step 1: Collect sentiment data: Use web crawler technology to automatically capture customer sentiment data from multiple social platforms, online insurance review areas, and customer service interaction records, clean and pre-process the collected customer sentiment data, and obtain a sentiment data set for subsequent sentiment analysis; Step 2: Conduct sentiment analysis, using deep learning algorithms, combined with professional sentiment vocabulary and weight settings, to accurately identify and quantify customer sentiment expressions, and accurately convert sentiment tendencies into risk preference levels; Step 3: Incorporate factors such as customer risk preferences obtained through sentiment analysis into the traditional insurance pricing model, use mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, build a dynamic pricing model, and calculate and adjust the customer's policy terms and premiums in real time.

2. The insurance pricing method based on web crawler technology according to claim 1, characterized in that: In step 1, the collected customer sentiment data is cleaned and preprocessed to obtain a sentiment data set for subsequent sentiment analysis. The method is as follows: Assume that the original collected text data set is D = {d1, d2, ..., d n }, where d i Represents the i-th text data. These text data are segmented during the data collection phase to build a high-quality data set. A large amount of insurance field text data is used to train the topic model to obtain the topic-word distribution. And the document-topic distribution φ, for the text to be segmented d i , in the traditional word segmentation results Based on the information of the topic model, we first find two adjacent words w ij and w i(j+1) , calculate the probability P(w ij , w i(j+1) |z), where z represents the topic; if the probability is greater than the empirical coefficient 0.73, the two words are merged into one phrase.

3. The insurance pricing method based on web crawler technology according to claim 1, characterized in that: In step 2, a deep learning algorithm is used, and a professional sentiment word library and weight setting are combined to achieve accurate recognition and quantitative analysis of customer sentiment expression. The method of accurately converting sentiment tendency into risk preference degree is to design a deep learning network that integrates attention mechanism, multimodal information and reinforcement learning feedback; the deep learning network is used for sentiment analysis and risk quantification judgment; The overall structure of the deep learning network is divided into an input layer, a feature extraction layer, an attention mechanism layer, a multimodal fusion layer, a risk quantification layer, and a reinforcement learning feedback layer; The input layer is used to receive text data, image data and voice data; Among them, text data: the pre-processed insurance-related text data is converted into word vector representation, and the word vector is obtained using Word2Vec word embedding technology to form a sequence input; Image data: The collected images such as insurance promotional posters and claims-related pictures are preprocessed to make them suitable for the input of convolutional neural networks. Voice data: The collected voice signals such as customer service call recordings and customer feedback audio are converted into a processable form of spectrogram through methods such as short-time Fourier transform.

4. The insurance pricing method based on web crawler technology according to claim 3 is characterized in that: The feature extraction layer is used to extract features from text data, image data, and voice data: The feature extraction layer extracts text features of text data by using a bidirectional long short-term memory network (Bi-LSTM) to process the text word vector sequence, which can simultaneously capture the forward and backward information of the text, effectively handle the long sequence dependency problem, and learn to obtain rich text semantic features; The feature extraction layer extracts features of the image data by using a convolutional neural network ResNet50 to extract features of the image data, and learns common image features on a large-scale image dataset to capture key information related to insurance in the image; The feature extraction layer extracts the speech features of the speech data by using a convolutional neural network to perform a convolution operation on the spectrum graph to extract the acoustic features of the speech.

5. The insurance pricing method based on web crawler technology according to claim 3 is characterized in that: The attention mechanism layer applies attention mechanism to the feature representation of text, image and speech respectively, and highlights the text parts that are more important for sentiment analysis and risk quantification by calculating the attention weights between the hidden states at each time step; similarly, the image features and speech features are also processed by attention mechanism to focus on key information; The detailed process is as follows: Three different features are used: text features text features Where T is the length of the text sequence, d T It is the text feature dimension; image feature Where N is the number of image features, d I is the image feature dimension; voice feature Where M is the number of speech features, d S is the speech feature dimension; the above-mentioned features are obtained after being processed by the feature extraction layer of each modality. During feature transformation, we perform linear transformation on the features of each modality to make them of the same dimension for subsequent calculation; for text features T, we use the linear transformation matrix W T Until T′=TW T ,in d is the unit after unified measurement; for image feature I, through the linear transformation matrix W I We get I′=IW I ,in For the speech feature S, through the linear transformation matrix W S , we get S′=SW S ,in By calculating the attention scores of these three features and normalizing each attention score matrix using the Softmax function, we can get the attention weight matrix: Text-image attention weight matrix W TI : Text-speech attention weight matrix W TS : Image-speech attention weight matrix W IS : After the above cross-modal attention mechanism is processed, the text feature T that integrates different modal information is obtained. fused 、Image Features I fused and speech feature S fused The fused features will be used as the input of the subsequent multimodal fusion layer, and further integrated to generate a unified feature representation containing rich cross-modal correlation information.

6. The insurance pricing method based on web crawler technology according to claim 3 is characterized in that: The risk quantification layer is used to perform nonlinear transformation on the basis of fusion features using a multi-layer perceptron to map the fusion features to the prediction space of sentiment scores and risk preferences; the sentiment scores are converted into probability distributions of different sentiment categories (positive, negative, neutral) through the Softmax layer, and the quantitative values ​​of risk preferences are output at the same time.

7. The insurance pricing method based on web crawler technology according to claim 1, characterized in that: The specific process of integrating factors such as customer risk preferences obtained through sentiment analysis into the traditional insurance pricing model, using mathematical modeling and statistical analysis methods to determine the weight of each factor in pricing, building a dynamic pricing model, and calculating and adjusting the customer's policy terms and premiums in real time is as follows: The risk preference level (R) of the customer is obtained through sentiment analysis. The value range is between ([0, 1]), (0) indicates a very high risk aversion level, and (1) indicates a very high risk preference level. In order to reflect the impact of sentiment analysis factors on pricing, a dynamic adjustment coefficient (k) is introduced; (k=αR+β) Among them, (α) and (β) are parameters determined through a large amount of historical data and statistical analysis. Dynamic Pricing (P dynamic ) is: (P dynamic =P traditional ×(1+k)) Taking into account dynamic factors such as market competition and customer feedback, a market dynamic factor (M) is introduced, which is determined by factors such as market supply and demand, competitor pricing, etc., and its value range is between ([0.8, 1.2]). The final dynamic pricing formula is: (P dynamic =P traditional ×(1+k)×M) For product recommendations, we built a personalized insurance recommendation system based on customers' emotional data and behavioral data (combined with emotional factors) to recommend insurance products that best meet customers' needs and expectations.

8. The insurance pricing method based on web crawler technology according to claim 7, characterized in that: For the product recommendation part, we built a personalized insurance recommendation system based on the customer's emotional data and behavioral data, combined with emotional factors. The specific process of recommending the insurance products that best meet the customer's needs and expectations is as follows: First, the emotional tendency vector (E i ) (e.g., numerical representations of emotional dimensions such as positive, negative, and neutral), each insurance product \(j\) also has a corresponding emotional adaptation vector (F j ) (determined according to product characteristics and target customers’ emotional needs); Then calculate the matching degree between customer sentiment and product sentiment (match(i, j)): Next, the sentiment matching degree is integrated into the prediction score to obtain the final recommendation score (S ij ): Among them, (ω∈[0,1]) is a weight coefficient used to balance the impact of the prediction score based on collaborative filtering and the sentiment matching degree. According to (S ij ) Sort products and recommend products with higher ratings to customers.