Insurance risk prediction method and device, equipment and medium
By acquiring multi-source data through intelligent agents, transforming and extracting strategies to form a business data warehouse, and combining it with multi-layer neural networks for insurance risk prediction, the problem of low prediction efficiency in existing technologies has been solved, enabling accurate identification of life insurance fraud and reduction of economic losses.
Patent Information
- Application Number
- CN202511207719.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-12-12
AI Technical Summary
Existing methods for predicting insurance risks lack the ability to deeply mine and analyze correlations, making it difficult to identify fraud patterns and relationships in complex data, resulting in low prediction efficiency.
By acquiring insurance applications, multiple data collection agents are used to collect raw datasets from multiple sources. A business data warehouse is formed based on a transformation strategy. A feature matrix is obtained using an extraction strategy, and risk prediction is performed using a prediction strategy. Fraud behavior is identified by combining a multi-layer neural network.
It has enabled accurate identification of life insurance fraud, reduced the rate of false positives and false negatives, reduced the economic losses of insurance companies caused by fraud, and improved prediction efficiency.
Smart Images

Figure CN121120267A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a method and device for predicting insurance risks, and a medium. BACKGROUND
[0002] Currently, the prior art lacks the ability to deeply mine and analyze complex data, and especially as insurance fraud methods become increasingly covert and diverse, a single analysis method is difficult to identify fraud patterns and correlation relationships hidden in the data. For example, it is difficult to effectively discover the implicit correlation between different insurance applicants, the difference between false health certificates and actual health conditions, and the like, making it difficult to detect fraud in a timely manner, thereby resulting in low prediction efficiency.
[0003] Therefore, the existing method for predicting insurance risks has the problem of low prediction efficiency. SUMMARY
[0004] Embodiments of the present application provide a method and device for predicting insurance risks, and a medium, aiming to solve the problem of low prediction efficiency of the prior art method for predicting insurance risks.
[0005] To solve the above problems, in a first aspect, the embodiments of the present application provide a method for predicting insurance risks, comprising:
[0006] obtaining an insurance application;
[0007] using the insurance application as a collection condition and collecting a plurality of original data sets by a plurality of collection agents;
[0008] processing the plurality of original data sets based on a conversion strategy to obtain a business data warehouse;
[0009] processing the business data warehouse based on an extraction strategy to obtain a feature matrix;
[0010] processing the feature matrix based on a prediction strategy to obtain a prediction result.
[0011] In a second aspect, the embodiments of the present application provide a device for predicting insurance risks, comprising:
[0012] an obtaining unit configured to obtain an insurance application;
[0013] a collection unit configured to use the insurance application as a collection condition and collect a plurality of original data sets by a plurality of collection agents;
[0014] a conversion unit configured to process the plurality of original data sets based on a conversion strategy to obtain a business data warehouse;
[0015] an extraction unit configured to perform extraction processing on the service data warehouse according to an extraction strategy to obtain a feature matrix;
[0016] a prediction unit configured to perform prediction processing on the feature matrix according to a prediction strategy to obtain a prediction result.
[0017] In a third aspect, an embodiment of the present application provides a computer device, which comprises a memory and a processor connected to the memory; the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to perform the method in the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a storage medium, which stores a computer program, and the computer program comprises program instructions, which, when executed by a processor, implement the method in the first aspect.
[0019] The embodiments of the present application provide a risk prediction method, device, equipment and medium for insurance application, and the method comprises the following steps: obtaining an insurance application; taking the insurance application as a collection condition, and collecting a multi-source original data set by a plurality of collection agents; performing conversion processing on the multi-source original data set based on a conversion strategy to obtain a service data warehouse; performing extraction processing on the service data warehouse based on an extraction strategy to obtain a feature matrix; and performing prediction processing on the feature matrix based on a prediction strategy to obtain a prediction result. Therefore, the embodiments of the present application take the insurance application as the collection condition, collect the multi-source original data set by the plurality of collection agents, perform the conversion processing on the multi-source original data set based on the conversion strategy to obtain the service data warehouse, perform the extraction processing on the service data warehouse based on the extraction strategy to obtain the feature matrix, and perform the prediction processing on the feature matrix based on the prediction strategy to obtain the prediction result, so that the plurality of collection agents are used to ensure the comprehensiveness of the multi-source data, the feature extraction and the prediction strategy are combined, various hidden life insurance fraud behaviors can be accurately identified, the misjudgment and omission rates are greatly reduced, the economic losses of insurance companies caused by fraud are effectively reduced, and the prediction efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0021] Figure 1 A flowchart of the risk prediction method for insurance application provided by the embodiments of the present application is shown in the figure.
[0022] Figure 2A schematic block diagram of the insurance risk prediction device provided by the embodiment of the present application is shown in the figure.
[0023] Figure 3 A schematic block diagram of the computer device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0025] It should be understood that, when used in the specification and the appended claims, the terms "comprise" and "include" indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0026] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] It should be further understood that the term "and / or" used in the specification and the appended claims of the present application means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0028] Please refer to Figure 1 , Figure 1 A flowchart of the insurance risk prediction method provided by the embodiment of the present application is shown in the figure. As Figure 1 shown, the present application provides an insurance risk prediction method, which comprises the following steps S110-S150.
[0029] S110, obtaining an insurance application.
[0030] In the present embodiment, the application scenario of the present solution can be the financial field and the like, and specifically, can be applied to the insurance business field such as life insurance and the like.
[0031] The insurance application can include applications of various types of insurance such as life insurance and the like.
[0032] In an embodiment, before the step of obtaining the insurance application, the method further comprises:
[0033] According to the data scene, a plurality of collection types are obtained through division processing;
[0034] The plurality of collection types are used for construction processing to obtain the plurality of collection agents.
[0035] In the embodiment, the plurality of collection types are obtained through division processing according to the data scene, specifically, the data scene can include an applicant basic information data scene, an insurance application history data scene, a claim data scene, an external environment data scene, etc.; each data scene corresponds to a collection type; and the plurality of collection types can include applicant basic information data collection, insurance application history data collection, claim data collection, external environment data collection, etc.
[0036] The plurality of collection types are used for construction processing to obtain the plurality of collection agents, specifically, the plurality of collection agents are obtained through construction processing according to each collection type in the plurality of collection types. One collection type corresponds to one collection agent, that is, the plurality of collection agents can include an applicant basic information data collection agent, an insurance application history data collection agent, a claim data collection agent, an external environment data collection agent, etc.
[0037] The applicant basic information collection agent collects the data of the identity information, health records, credit records, and income proofs of the applicant by connecting with the social security system, medical institutions, credit investigation agencies, etc.; the insurance application history data collection agent obtains the information of the past insurance application records, insurance cancellation records, and insurance amount changes of the applicant from the internal database of the insurance company; the claim data collection agent collects the detailed information of the historical claim cases, including the claim reason, claim amount, claim time, audit result, etc.; and the external environment data collection agent pays attention to the data of the fraud cases in the industry, changes in regulatory policies, and the influence of the economic environment on the life insurance fraud, etc.
[0038] Through the above embodiment, it can be known that the insurance application is obtained, and before the insurance application is obtained, the plurality of collection types are obtained through division processing according to the data scene, and the plurality of collection agents are obtained through construction processing using the plurality of collection types. Therefore, the plurality of collection agents are obtained through division and construction processing according to the data scene and the plurality of collection types, the comprehensiveness of data collection is ensured, the prediction efficiency is improved, and subsequent targeted processing is performed according to the insurance application and the plurality of collection agents.
[0039] S120, taking the insurance application as a collection condition, and collecting a plurality of original data sets by a plurality of collection agents.
[0040] In the embodiment, after the insurance application is obtained and the plurality of collection agents are constructed, the insurance application is taken as a collection condition, and a plurality of collection agents are used to collect a multi-source original data set. The insurance application can include name, ID number, insurance type, etc.
[0041] The plurality of collection agents collect data according to the collection condition of the insurance application, and after the collected data is preliminarily screened and standardized, the collected data is output to a data aggregation node to form the multi-source original data set.
[0042] Through the above embodiment, it can be known that the insurance application is taken as a collection condition, and a plurality of collection agents are used to collect a multi-source original data set, which ensures the comprehensiveness of data collection, improves the prediction efficiency, and performs subsequent targeted processing according to the multi-source original data set.
[0043] S130, converting the multi-source original data set based on a conversion strategy to obtain a business data warehouse.
[0044] In the embodiment, after the multi-source original data set is determined, the multi-source original data set is converted based on a conversion strategy to obtain the business data warehouse.
[0045] In an embodiment, the multi-source original data set is converted based on a conversion strategy to obtain a business data warehouse, including:
[0046] The multi-source original data set is analyzed and processed to obtain abnormal data.
[0047] The abnormal data is cleaned based on a data cleaning strategy to obtain a multi-source intermediate data set.
[0048] The multi-source intermediate data set is input into a preset model to transform the multi-source intermediate data set to obtain a data vector.
[0049] The data vector is fused based on a data fusion algorithm to obtain the business data warehouse.
[0050] In the embodiment, the multi-source original data set is analyzed and processed to obtain abnormal data, specifically, the multi-source original data set is analyzed and processed to obtain missing values, abnormal values and repeated values; the missing values, the abnormal values and the repeated values are taken as the abnormal data.
[0051] The data cleaning strategy is used to clean the abnormal data to obtain a multi-source intermediate data set. Specifically, when the abnormal data is a missing value, the mean filling method is used to process the missing value; when the abnormal data is an abnormal value, the 3σ principle is used to identify and eliminate the abnormal value; and when the abnormal data is a repeated value, the repeated value is directly eliminated.
[0052] The multi-source intermediate data set is input into a preset model to transform the multi-source intermediate data set to obtain a data vector. Specifically, after the multi-source intermediate data set is input into the preset model, it is assumed that a text set corresponding to the multi-source intermediate data set can be represented as T=\{t_1,t_2,...,t_m\}, a vocabulary can be represented as V=\{v_1,v_2,...,v_n\}, and a bag-of-words vector of a text t_i can be represented as X_i=\(x_{i1},x_{i2},...,x_{in}\), where x_{ij} is the number of occurrences of the vocabulary v_j in the text t_i. The bag-of-words vector of each text in the text set is used as the data vector. The preset model can be a bag-of-words model.
[0053] The data fusion algorithm is used to fuse the data vector to obtain the business data warehouse. Specifically, the data fusion algorithm is used to integrate different sources and different types of data vectors to construct a unified business data warehouse, thereby providing a high-quality data basis for subsequent fraud risk analysis.
[0054] According to the above embodiments, the multi-source original data set is analyzed and processed to obtain abnormal data, the data cleaning strategy is used to clean the abnormal data to obtain a multi-source intermediate data set, the multi-source intermediate data set is input into a preset model to transform the multi-source intermediate data set to obtain a data vector, and the data fusion algorithm is used to fuse the data vector to obtain the business data warehouse. Therefore, the multi-source original data set is transformed based on the transformation strategy to obtain the business data warehouse, ensuring the comprehensiveness and accuracy of the data to improve the prediction efficiency, and subsequent targeted processing is performed based on the business data warehouse.
[0055] S140, a feature matrix is obtained by extracting the business data warehouse using an extraction strategy.
[0056] In this embodiment, after obtaining the business data warehouse, the extraction strategy is used to extract the business data warehouse to obtain a feature matrix.
[0057] In an embodiment, the extraction strategy is used to extract the business data warehouse to obtain a feature matrix, including:
[0058] obtain an original data matrix of the service data warehouse, and calculate a mean vector of the original data matrix;
[0059] perform first calculation processing on the original data matrix and the mean vector by using a preset formula to obtain a covariance matrix;
[0060] perform second calculation processing on the covariance matrix by using a decomposition algorithm to obtain a plurality of eigenvalues and a plurality of corresponding eigenvectors;
[0061] perform sorting processing on the plurality of eigenvalues and the plurality of corresponding eigenvectors in descending order to obtain an ordered feature set;
[0062] select eigenvectors corresponding to a preset number of eigenvalues in the ordered feature set to form a projection matrix;
[0063] perform third calculation processing on the original data matrix and the projection matrix to obtain the feature matrix.
[0064] In this embodiment, the original data matrix of the service data warehouse is assumed to be X_{n\times p} (n is the number of samples, and p is the number of features), and the mean vector of the original data matrix is obtained by calculation.
[0065] The first calculation processing on the original data matrix and the mean vector by using the preset formula to obtain the covariance matrix C is specifically performed by using the preset formula C=\frac{1}{n-1}X^T(X-\bar{X}).
[0066] The second calculation processing on the covariance matrix by using the decomposition algorithm to obtain the plurality of eigenvalues and the plurality of corresponding eigenvectors is specifically performed by using a characteristic equation C\alpha=\lambda\alpha to perform calculation processing on the covariance matrix to obtain the plurality of eigenvalues \lambda and the plurality of corresponding eigenvectors \alpha. The plurality of eigenvalues and the plurality of corresponding eigenvectors are in one-to-one correspondence, that is, one eigenvalue corresponds to one eigenvector.
[0067] After the sorting processing on the plurality of eigenvalues and the plurality of corresponding eigenvectors in descending order to obtain the ordered feature set, the eigenvectors corresponding to the preset number of eigenvalues in the ordered feature set are selected to form the projection matrix Z.
[0068] The third calculation processing is performed on the original data matrix and the projection matrix to obtain the feature matrix, specifically, the original data matrix and the projection matrix are multiplied to obtain the feature matrix, that is, the feature matrix after dimension reduction is Y=XZ.
[0069] According to the above embodiment, the original data matrix of the business data warehouse is obtained, and the mean vector of the original data matrix is calculated. The original data matrix and the mean vector are processed by a first calculation according to a preset formula to obtain a covariance matrix. A decomposition algorithm is used to perform a second calculation on the covariance matrix to obtain a plurality of eigenvalues and a plurality of corresponding eigenvectors. The plurality of eigenvalues and the plurality of corresponding eigenvectors are sorted in descending order to obtain an ordered feature set. The eigenvectors corresponding to the first preset number of eigenvalues in the ordered feature set are selected to form a projection matrix. The original data matrix and the projection matrix are processed by a third calculation to obtain the feature matrix. Therefore, the feature matrix is obtained by extracting the business data warehouse using the extraction strategy, which can accurately identify various hidden life insurance fraud behaviors, greatly reduce the misjudgment and omission rate, effectively reduce the economic losses caused by fraud for insurance companies, and improve the prediction efficiency.
[0070] S150, according to the prediction strategy, the feature matrix is processed to obtain a prediction result.
[0071] In this embodiment, after obtaining the feature matrix, the feature matrix is processed according to the prediction strategy to obtain a prediction result.
[0072] In one embodiment, the feature matrix is processed according to the prediction strategy to obtain a prediction result, comprising:
[0073] The feature matrix is input into a multi-layer neural network;
[0074] The number of layers of the multi-layer neural network and the weight matrix and bias term corresponding to each layer are obtained;
[0075] The weight matrix and bias term corresponding to each layer are calculated respectively to obtain the output value of each layer, and the output values of each layer are combined as an output value set;
[0076] The output value set is calculated using an activation function to obtain a risk probability;
[0077] The prediction result is determined according to the risk probability.
[0078] In one embodiment, the prediction result is determined according to the risk probability, comprising:
[0079] determining that the prediction result is high risk when the risk probability is greater than a risk threshold value;
[0080] determining that the prediction result is low risk when the risk probability is less than or equal to a risk threshold value.
[0081] In the embodiment, the feature matrix is input into a multi-layer neural network, assuming that the feature matrix x = (x_1, x_2, …, x_k); the number of layers L of the multi-layer neural network and the weight matrix W_l and the bias term b_l corresponding to each layer are obtained; the output value of each layer is obtained by calculating and processing the weight matrix and the bias term corresponding to each layer respectively, and the output values of each layer are combined as an output value set, and the risk probability is obtained by calculating and processing the output value set using an activation function, specifically by calculating formula P = \sigma(W_{out}h_{L}+b_{out}) = \frac{1}{1+e^{-(W_{out}h_{L}+b_{out})}}, L is the number of hidden layers, W_{out} and b_{out} are the weight and bias of the output layer respectively, and P is the risk probability. The activation function can be a ReLU function.
[0082] When the risk probability is greater than a risk threshold value, the prediction result is marked as high risk, i.e. high risk of fraud suspicion; when the risk probability is less than or equal to a risk threshold value, the prediction result is marked as low risk, i.e. low risk of fraud suspicion.
[0083] Through the above embodiment, it can be known that the feature matrix is input into a multi-layer neural network; the number of layers L of the multi-layer neural network and the weight matrix and the bias term corresponding to each layer are obtained; the output value of each layer is obtained by calculating and processing the weight matrix and the bias term corresponding to each layer respectively, and the output values of each layer are combined as an output value set; the risk probability is obtained by calculating and processing the output value set using an activation function; and the prediction result is determined according to the risk probability. Therefore, the use of a plurality of collection agents ensures the comprehensiveness of multi-source data, and the combination of feature extraction and prediction strategies can accurately identify various hidden life insurance fraud behaviors, greatly reduce the misjudgment and omission rate, effectively reduce the economic losses of insurance companies caused by fraud, and thus improve the prediction efficiency.
[0084] In an embodiment, after the prediction result is determined according to the risk probability, the method further comprises:
[0085] When the prediction result is high risk, analyzing and processing the insurance application and the multi-source original data set to obtain risk features;
[0086] outputting the risk features to obtain a risk assessment report.
[0087] In the embodiment, when the risk probability is greater than the risk threshold, the prediction result is marked as high risk, the risk features are obtained by analyzing and processing the insurance application and the multi-source original data set, and the risk assessment report is output according to the risk features.
[0088] The scheme can further include a decision-making agent and a feedback agent. The prediction result and the corresponding risk assessment report are input to the decision-making agent, and specific anti-fraud measures are formulated according to preset anti-fraud rules and strategies. For example, for an insurance application with high risk of fraud suspicion, the decision-making agent issues an instruction to further verify the information and requires the policyholder to supplement relevant proof materials. For a recognized fraud claim case, the decision-making agent issues an instruction to reject the claim and start an investigation. At the same time, the feedback agent tracks and collects the results of decision execution, including verified information, investigation results, final confirmation of fraud cases, and other data, and feeds these data back to the data preprocessing and fusion agent and the risk analysis and prediction agent. The data preprocessing and fusion agent updates the data warehouse using the feedback data, and the risk analysis and prediction agent adjusts the model parameters, optimizes the feature weights and prediction algorithms, and continuously improves the accuracy and effectiveness of anti-fraud risk control and prediction, forming a closed-loop optimization system.
[0089] Through the above embodiments, when the prediction result is high risk, the risk features are obtained by analyzing and processing the insurance application and the multi-source original data set; and the risk assessment report is output according to the risk features. Therefore, the prediction result and the risk assessment report are used to predict potential fraud risks in advance, so that the insurance company can take effective preventive measures before fraud occurs, change from passive response to active prevention, and control risks in the embryonic state. Through continuous optimization of the feedback mechanism, the system can quickly adapt to changes in life insurance fraud methods and new fraud patterns, timely adjust anti-fraud strategies and model parameters, and maintain long-term effective anti-fraud capabilities. The collaborative work mode of multiple agents improves the efficiency of anti-fraud work. Each agent processes tasks in parallel, reducing the time cost of data collection, analysis, and decision-making, so that the insurance company can quickly respond to anti-fraud needs in life insurance business, improve overall operational efficiency and service quality, and enhance market competitiveness.
[0090] In summary, this embodiment of the invention obtains insurance applications; uses these applications as collection conditions, and collects them through several intelligent collection agents to obtain a multi-source raw dataset; transforms the multi-source raw dataset based on a transformation strategy to obtain a business data warehouse; extracts features from the business data warehouse using an extraction strategy; and performs prediction processing on the feature matrix according to a prediction strategy to obtain a prediction result. Therefore, this embodiment of the invention, by using insurance applications as collection conditions and collecting them through several intelligent collection agents to obtain a multi-source raw dataset; transforming the multi-source raw dataset based on a transformation strategy to obtain a business data warehouse; extracting features from the business data warehouse using an extraction strategy; and performing prediction processing on the feature matrix according to a prediction strategy to obtain a prediction result, ensures the comprehensiveness of multi-source data by utilizing several intelligent collection agents. Combined with feature extraction and prediction strategies, it can accurately identify various hidden life insurance fraud behaviors, significantly reduce false positives and false negatives, effectively reduce the economic losses of insurance companies caused by fraud, and thus improve prediction efficiency.
[0091] Figure 2 This is a schematic block diagram of an insurance risk prediction device provided in an embodiment of the present invention. Figure 2 As shown, this embodiment of the invention provides an insurance risk prediction device 700 that implements the method described above. Specifically, please refer to... Figure 2 The insurance risk prediction device 700 includes:
[0092] Acquisition unit 701 is used to acquire insurance applications;
[0093] The acquisition unit 702 is used to take the insurance application as an acquisition condition and acquire it by a number of acquisition agents to obtain a multi-source raw dataset.
[0094] The transformation unit 703 is used to transform the multi-source original dataset based on the transformation strategy to obtain a business data warehouse;
[0095] Extraction unit 704 is used to extract features from the business data warehouse using an extraction strategy.
[0096] The prediction unit 705 is used to perform prediction processing on the feature matrix according to the prediction strategy to obtain the prediction result.
[0097] In some embodiments, when performing the step of transforming the multi-source raw dataset based on the transformation strategy to obtain a business data warehouse, the transformation unit 703 is specifically used for:
[0098] The abnormal data were obtained by analyzing and processing the multi-source raw dataset;
[0099] perform data cleaning processing on the abnormal data based on a data cleaning strategy to obtain a multi-source intermediate data set;
[0100] input the multi-source intermediate data set into a preset model to perform transformation processing on the multi-source intermediate data set to obtain a data vector;
[0101] perform fusion processing on the data vector based on a data fusion algorithm to obtain the business data warehouse.
[0102] In some embodiments, the extraction unit 704, when performing the processing step of extracting the business data warehouse based on the extraction strategy to obtain a feature matrix, is specifically configured to:
[0103] obtain an original data matrix of the business data warehouse and calculate a mean vector of the original data matrix;
[0104] perform first calculation processing on the original data matrix and the mean vector based on a preset formula to obtain a covariance matrix;
[0105] perform second calculation processing on the covariance matrix using a decomposition algorithm to obtain a plurality of eigenvalues and a plurality of corresponding eigenvectors;
[0106] perform sorting processing on the plurality of eigenvalues and the plurality of corresponding eigenvectors in descending order to obtain an ordered feature set;
[0107] select eigenvectors corresponding to a preset number of eigenvalues in the ordered feature set to form a projection matrix;
[0108] perform third calculation processing on the original data matrix and the projection matrix to obtain the feature matrix.
[0109] In some embodiments, the prediction unit 705, when performing the processing step of predicting the feature matrix based on the prediction strategy to obtain a prediction result, is specifically configured to:
[0110] input the feature matrix into a multi-layer neural network;
[0111] obtain the number of layers of the multi-layer neural network and the weight matrix and bias term corresponding to each layer;
[0112] perform calculation processing on each layer using the weight matrix and bias term corresponding to each layer to obtain the output value of each layer, and combine the output values of each layer as an output value set;
[0113] perform calculation processing on the output value set using an activation function to obtain a risk probability;
[0114] determine the prediction result according to the risk probability.
[0115] In some embodiments, the prediction unit 705, in the process of determining the prediction result according to the risk probability, is specifically configured to:
[0116] determining the prediction result as high risk when the risk probability is greater than a risk threshold value;
[0117] determining the prediction result as low risk when the risk probability is less than or equal to the risk threshold value.
[0118] In some embodiments, the prediction unit 705, after the process of determining the prediction result according to the risk probability, is further configured to:
[0119] when the prediction result is high risk, performing analysis and processing based on the insurance application and the multi-source original data set to obtain risk characteristics;
[0120] outputting according to the risk characteristics to obtain a risk assessment report.
[0121] In some embodiments, the acquisition unit 701, before the process of acquiring the insurance application, is further configured to:
[0122] performing division processing according to a data scenario to obtain a plurality of collection types;
[0123] performing construction processing using the plurality of collection types to obtain the plurality of collection agents.
[0124] It should be noted that the specific implementation process of the above device can be clearly understood by those skilled in the art, and the corresponding description in the foregoing method embodiments can be referred to. For the convenience and brevity of description, it will not be repeated here.
[0125] The above device can be implemented in the form of a computer program, which can run on a computer device such as the computer device shown in Figure 3 .
[0126] Please refer to Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the application. The electronic device 800 can be a terminal or a server, wherein the terminal can be an electronic device with communication function. The server can be a stand-alone server or a server cluster composed of multiple servers.
[0127] Referring to Figure 3 , the electronic device 800 includes a processor 802, a memory, and a network interface 805 connected through a system bus 801, wherein the memory can include a non-volatile storage medium 803 and an internal memory 804.
[0128] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions which, when executed, can cause the processor 802 to perform an insurance risk prediction method.
[0129] The processor 802 is configured to provide computing and control capabilities to support the operation of the entire electronic device 800.
[0130] The non-volatile storage medium 803 provides an environment for the computer program 8032 stored therein to be executed by the processor 802, which can cause the processor 802 to perform an insurance risk prediction method.
[0131] The network interface 805 is configured to communicate with other devices over a network. Those skilled in the art can understand that the network interface 805 can be configured to communicate with other devices over a network according to the structure shown in the figure, and does not constitute a limitation on the electronic device 800 to which the present application is applied. The specific electronic device 800 can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement. Figure 3
[0132] The processor 802 is configured to run the computer program 8032 stored in the memory to implement the following steps:
[0133] Obtain an insurance application;
[0134] The insurance application is used as a collection condition, and a plurality of collection agents are used to collect a multi-source original data set;
[0135] The multi-source original data set is processed based on a conversion strategy to obtain a business data warehouse;
[0136] The business data warehouse is processed using an extraction strategy to obtain a feature matrix;
[0137] The feature matrix is processed according to a prediction strategy to obtain a prediction result.
[0138] In some embodiments, when the processor 802 implements the processing step of processing the multi-source original data set based on the conversion strategy to obtain the business data warehouse, it is specifically configured to:
[0139] The multi-source original data set is analyzed and processed to obtain abnormal data;
[0140] The abnormal data is processed based on a data cleaning strategy to obtain a multi-source intermediate data set;
[0141] inputting the multi-source intermediate data set into a preset model to transform the multi-source intermediate data set to obtain a data vector;
[0142] performing fusion processing on the data vector based on a data fusion algorithm to obtain the business data warehouse.
[0143] In some embodiments, the processor 802, when implementing the processing step of extracting the business data warehouse using the extraction strategy to obtain a feature matrix, is specifically configured to:
[0144] obtain an original data matrix of the business data warehouse and calculate a mean vector of the original data matrix;
[0145] perform first calculation processing on the original data matrix and the mean vector using a preset formula to obtain a covariance matrix;
[0146] perform second calculation processing on the covariance matrix using a decomposition algorithm to obtain a plurality of eigenvalues and a plurality of corresponding eigenvectors;
[0147] perform sorting processing on the plurality of eigenvalues and the plurality of corresponding eigenvectors in descending order to obtain an ordered feature set;
[0148] select eigenvectors corresponding to a preset number of eigenvalues in the ordered feature set to form a projection matrix;
[0149] perform third calculation processing on the original data matrix and the projection matrix to obtain the feature matrix.
[0150] In some embodiments, the processor 802, when implementing the processing step of predicting the feature matrix according to a prediction strategy to obtain a prediction result, is specifically configured to:
[0151] input the feature matrix into a multi-layer neural network;
[0152] obtain the number of layers of the multi-layer neural network and the weight matrix and bias term corresponding to each layer;
[0153] perform calculation processing on each layer using the weight matrix and bias term corresponding to each layer to obtain the output value of each layer, and combine the output values of each layer as an output value set;
[0154] perform calculation processing on the output value set using an activation function to obtain a risk probability;
[0155] determine the prediction result according to the risk probability.
[0156] In some embodiments, the processor 802, when implementing the processing step of determining the prediction result according to the risk probability, is specifically configured to:
[0157] determining that the prediction result is high risk when the risk probability is greater than a risk threshold;
[0158] determining that the prediction result is low risk when the risk probability is less than or equal to a risk threshold.
[0159] In some embodiments, the processor 802, after implementing the processing step of determining the prediction result according to the risk probability, is further configured to:
[0160] performing analysis processing based on the insurance application and the multi-source original data set to obtain risk features when the prediction result is high risk;
[0161] outputting a risk assessment report according to the risk features.
[0162] In some embodiments, the processor 802, before implementing the processing step of obtaining the insurance application, is further configured to:
[0163] performing division processing according to data scenarios to obtain a plurality of collection types;
[0164] performing construction processing using the plurality of collection types to obtain the plurality of collection agents.
[0165] It should be understood that, in the embodiments of the present application, the processor 802 can be a central processing unit (CPU), and the processor 802 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0166] Those of ordinary skill in the art can understand that all or part of the processes in the method of the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments of the method.
[0167] Therefore, the application further provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by a processor to cause the processor to perform the following steps:
[0168] obtaining an insurance application;
[0169] taking the insurance application as a collection condition, and collecting a multi-source original data set by a plurality of collection agents;
[0170] performing conversion processing on the multi-source original data set based on a conversion strategy to obtain a business data warehouse;
[0171] performing extraction processing on the business data warehouse based on an extraction strategy to obtain a feature matrix;
[0172] performing prediction processing on the feature matrix based on a prediction strategy to obtain a prediction result.
[0173] In an embodiment, when the processor executes the program instructions to implement the processing step of performing conversion processing on the multi-source original data set based on the conversion strategy to obtain the business data warehouse, the processor is specifically configured to:
[0174] performing analysis processing on the multi-source original data set to obtain abnormal data;
[0175] performing data cleaning processing on the abnormal data based on a data cleaning strategy to obtain a multi-source intermediate data set;
[0176] inputting the multi-source intermediate data set into a preset model to perform transformation processing on the multi-source intermediate data set to obtain a data vector;
[0177] performing fusion processing on the data vector based on a data fusion algorithm to obtain the business data warehouse.
[0178] In an embodiment, when the processor executes the program instructions to implement the processing step of performing extraction processing on the business data warehouse based on the extraction strategy to obtain the feature matrix, the processor is specifically configured to:
[0179] obtaining an original data matrix of the business data warehouse, and calculating a mean vector of the original data matrix;
[0180] performing first calculation processing on the original data matrix and the mean vector by using a preset formula to obtain a covariance matrix;
[0181] performing second calculation processing on the covariance matrix by using a decomposition algorithm to obtain a plurality of eigenvalues and a plurality of corresponding eigenvectors;
[0182] sort the plurality of characteristic values and the plurality of corresponding eigenvectors in descending order to obtain an ordered eigenvector set;
[0183] select eigenvectors corresponding to a preset number of characteristic values in the ordered eigenvector set to form a projection matrix;
[0184] perform third calculation processing on the original data matrix and the projection matrix to obtain the characteristic matrix.
[0185] In an embodiment, when the processor executes the program instructions to implement the processing step of performing prediction processing on the characteristic matrix according to a prediction strategy to obtain a prediction result, the processor is specifically configured to:
[0186] input the characteristic matrix into a multi-layer neural network;
[0187] obtain the number of layers of the multi-layer neural network and the weight matrix and bias term corresponding to each layer;
[0188] perform calculation processing on each layer using the weight matrix and bias term corresponding to the layer to obtain an output value of each layer, and combine the output values of each layer as an output value set;
[0189] perform calculation processing on the output value set using an activation function to obtain a risk probability;
[0190] determine the prediction result according to the risk probability.
[0191] In an embodiment, when the processor executes the program instructions to implement the processing step of determining the prediction result according to the risk probability, the processor is specifically configured to:
[0192] when the risk probability is greater than a risk threshold, determine that the prediction result is high risk;
[0193] when the risk probability is less than or equal to the risk threshold, determine that the prediction result is low risk.
[0194] In an embodiment, after the processor executes the program instructions to implement the processing step of determining the prediction result according to the risk probability, the processor is further configured to:
[0195] when the prediction result is high risk, perform analysis processing based on the insurance application and the multi-source original data set to obtain a risk feature;
[0196] output a risk assessment report according to the risk feature.
[0197] In an embodiment, before the processor executes the program instructions to implement the processing step of obtaining an insurance application, the processor is further configured to:
[0198] According to the data scene, a plurality of collection types are obtained through division processing;
[0199] The plurality of collection types are obtained through construction processing to obtain the plurality of collection agents.
[0200] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, and various computer-readable storage media that can store program codes.
[0201] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0202] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of each unit is only a logical functional division, and actual implementation can have another division method. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0203] The steps in the method embodiments of the present application can be adjusted, combined, and reduced in sequence according to actual needs. The units in the device embodiments of the present application can be combined, divided, and reduced according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0204] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make an electronic device (which can be a personal computer, a terminal, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application.
[0205] The above merely describes specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be encompassed within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims. The software tools, models or components appearing in the embodiments of the present application are merely illustrative and do not represent actual use.
Claims
1. A method for predicting insurance risks, characterized in that, The method comprises the following steps: obtaining an insurance application; collecting the insurance application as a collection condition, and collecting a multi-source original data set by a plurality of collection agents; processing the multi-source original data set based on a conversion strategy to obtain a business data warehouse; processing the business data warehouse based on an extraction strategy to obtain a feature matrix; processing the feature matrix based on a prediction strategy to obtain a prediction result.
2. The method of claim 1, wherein, The method of processing the multi-source original data set based on the conversion strategy to obtain the business data warehouse comprises the following steps: analyzing and processing the multi-source original data set to obtain abnormal data; processing the abnormal data based on a data cleaning strategy to obtain a multi-source intermediate data set; inputting the multi-source intermediate data set into a preset model to transform the multi-source intermediate data set into a data vector; processing the data vector based on a data fusion algorithm to obtain the business data warehouse.
3. The method of claim 1, wherein, The method of processing the business data warehouse based on the extraction strategy to obtain the feature matrix comprises the following steps: obtaining an original data matrix of the business data warehouse and calculating a mean vector of the original data matrix; performing first calculation processing on the original data matrix and the mean vector based on a preset formula to obtain a covariance matrix; performing second calculation processing on the covariance matrix based on a decomposition algorithm to obtain a plurality of characteristic values and a plurality of corresponding characteristic vectors; performing sorting processing on the plurality of characteristic values and the plurality of corresponding characteristic vectors in descending order to obtain an ordered feature set; selecting characteristic vectors corresponding to a preset number of characteristic values in the ordered feature set to form a projection matrix; performing third calculation processing on the original data matrix and the projection matrix to obtain the feature matrix.
4. The method of claim 1, wherein, The method of processing the feature matrix based on the prediction strategy to obtain the prediction result comprises the following steps: inputting the feature matrix into a multi-layer neural network; obtaining the number of layers of the multi-layer neural network and the weight matrix and bias term corresponding to each layer; performing calculation processing on each layer based on the weight matrix and bias term corresponding to each layer to obtain the output value of each layer, and combining the output values of each layer as an output value set; performing calculation processing on the output value set based on an activation function to obtain a risk probability; determining the prediction result based on the risk probability.
5. The method of claim 1, wherein, The method of determining the prediction result based on the risk probability comprises the following steps: when the risk probability is greater than a risk threshold, determining that the prediction result is high risk; when the risk probability is less than or equal to the risk threshold, determining that the prediction result is low risk.
6. The method of claim 5, wherein, After determining the prediction result based on the risk probability, the method further comprises the following steps: when the prediction result is high risk, analyzing and processing the insurance application and the multi-source original data set to obtain a risk feature; outputting the risk feature to obtain a risk assessment report.
7. The method of claim 1, wherein, Before obtaining the insurance application, the method further comprises the following steps: dividing based on a data scene to obtain a plurality of collection types; constructing the plurality of collection types to obtain the plurality of collection agents.
8. An underwriting risk prediction apparatus, characterized by, The method comprises the following steps: an obtaining unit is configured to obtain an insurance application; The collection unit is configured to take the insurance application as a collection condition, and a plurality of collection agents are configured to collect a multi-source original data set; The conversion unit is configured to convert and process the multi-source original data set based on a conversion strategy to obtain a business data warehouse; The extraction unit is configured to extract and process the business data warehouse based on an extraction strategy to obtain a feature matrix; The prediction unit is configured to predict and process the feature matrix based on a prediction strategy to obtain a prediction result.
9. A computer device, comprising: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method in any one of claims 1-7 when executing the computer program.
10. A storage medium, characterized by The storage medium stores a computer program, the computer program includes program instructions, and the program instructions can implement the method in any one of claims 1-7 when executed by a processor.