Insurance anti-fraud method and device, electronic equipment and storage medium

By acquiring and analyzing multiple information from the target customers, and using the pre-trained fraud risk identification model for multi-dimensional analysis, the final risk score is calculated, and the problem of low accuracy in detecting insurance fraud in the prior art is solved, achieving higher accuracy and reliability.

CN120106990APending Publication Date: 2025-06-06CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510271897.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the detection of insurance fraud in the prior art, there are problems of inconsistent judgment results and emotional bias, resulting in a low accuracy rate.

Method used

By obtaining the target customer's customer information, current transaction information and historical transaction information, combining the mapping relationship between the pre-determined customer type and risk score, a pre-trained fraud risk identification model is used to perform multi-dimensional analysis to calculate the final risk score.

Benefits of technology

Through multi-dimensional probability calculation, false positives and missed reports are reduced, the accuracy of insurance anti-fraud is improved, and the limitations of a single dimension are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106990A_ABST
    Figure CN120106990A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning and the field of financial science and technology, and relates to an insurance anti-fraud method and device, electronic equipment and a storage medium. When a request for detecting the insurance fraud risk of a target customer sent by a client is received, obtaining customer information, current transaction information and historical transaction information of the target customer, determining a customer type according to the customer information, and determining a first risk score according to a mapping relationship between the customer type and the risk score; analyzing the current transaction information of the target customer by using the first fraud risk identification model to obtain a second risk score; analyzing the historical transaction information by using a second fraud risk identification model to obtain a third risk score; and finally, carrying out weighted summation on the three risk scores to obtain a final risk score, and feeding back the final risk score to the client. According to the invention, through multi-dimensional probability calculation, limitation of a single dimension is avoided, and misinformation and missed information can be reduced, so that the accuracy of insurance anti-fraud is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine learning technology and financial technology, and to an insurance anti-fraud method, device, electronic device and storage medium. Background Art

[0002] In modern society, financial insurance has become an indispensable part of people's lives, providing individuals and enterprises with an effective risk management tool. However, with the rapid development of the financial insurance market, insurance fraud has also increased, causing huge economic losses to insurance companies and damaging the interests of legitimate policyholders. Therefore, insurance anti-fraud has become an important measure to ensure financial security and market order.

[0003] In the existing field, in order to protect the interests of insurance companies and legitimate policyholders, testers are often required to conduct a detailed analysis of each insurance behavior. However, different personnel have different experience and knowledge levels, which may lead to different judgment results for the same case. In addition, during the manual detection process, the emotions and prejudices of the testers may affect the objectivity of the judgment. For example, preconceived views of certain customers may lead to misjudgment.

[0004] Therefore, how to improve the accuracy of detecting insurance fraud is an urgent problem to be solved. Summary of the invention

[0005] In view of the above, it is necessary to provide an insurance anti-fraud method to avoid the limitations of a single dimension, help reduce false positives and false negatives, and thus improve the accuracy of insurance anti-fraud.

[0006] To achieve the above object, the present invention provides an insurance anti-fraud method, characterized in that the method comprises:

[0007] After receiving a request from a client to detect insurance fraud risk of a target customer, obtaining relevant information of the target customer, the relevant information including customer information, current transaction information and historical transaction information;

[0008] Analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score;

[0009] Analyzing the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score;

[0010] Calculate a third risk score for the target customer based on the historical transaction information and the pre-trained second fraud risk identification model;

[0011] The first risk score, the second risk score and the third risk score are weightedly summed to obtain a final risk score of the target customer and feed it back to the client.

[0012] In addition, to achieve the above-mentioned purpose, the present invention also provides an insurance anti-fraud device, characterized in that the device comprises:

[0013] Relevant information acquisition module: used to acquire relevant information of a target customer after receiving a request sent by a client to detect the insurance fraud risk of the target customer, wherein the relevant information includes customer information, current transaction information and historical transaction information;

[0014] A first risk scoring module: used for analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score;

[0015] A second risk scoring module: used to analyze the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score;

[0016] A third risk scoring module: used for calculating a third risk score of the target customer based on the historical transaction information and the pre-trained second fraud risk identification model;

[0017] Risk score return module: used to perform weighted summation on the first risk score, the second risk score and the third risk score to obtain a final risk score of the target customer and feed it back to the client.

[0018] In addition, to achieve the above object, the present invention further provides an electronic device, the electronic device comprising:

[0019] a memory storing at least one computer program; and

[0020] The processor executes the program stored in the memory to implement the above-mentioned insurance anti-fraud method.

[0021] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned insurance anti-fraud method.

[0022] In the above technical solution provided by the present invention, by accepting a request to detect the insurance fraud risk of a target customer, the customer information, current transaction information and historical transaction information of the target customer are obtained, the customer type is determined according to the customer information, and the first risk score is determined according to the mapping relationship between the customer type and the risk score; the current transaction information of the target customer is analyzed using the first fraud risk identification model to obtain the second risk score; the historical transaction information is analyzed using the second fraud risk identification model to obtain the third risk score; finally, the three risk scores are weighted and summed to obtain the final risk score and fed back to the client. Through multi-dimensional probability calculation, the limitations of a single dimension are avoided, which helps to reduce false positives and false negatives, thereby improving the accuracy of insurance anti-fraud. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of an application environment of an insurance anti-fraud method in one embodiment of the present invention;

[0024] Figure 2 A schematic diagram of a flow chart of an insurance anti-fraud method provided by an embodiment of the present invention;

[0025] Figure 3 A schematic diagram of the structure of an insurance anti-fraud device provided by an embodiment of the present invention;

[0026] Figure 4 is a schematic diagram of a structure of a computer device in one embodiment of the present invention;

[0027] Figure 5 It is another structural schematic diagram of a computer device in one embodiment of the present invention.

[0028] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0030] It should be noted that the descriptions of "first", "second", etc. in the present invention are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0031] The insurance anti-fraud method provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the client communicates with the server through a network. After receiving the request sent by the client to detect the insurance fraud risk of the target customer, the server obtains the relevant information of the target customer, including customer information, current transaction information and historical transaction information; analyzes the customer type of the target customer according to the customer information, and determines the first risk score of the target customer according to the predetermined mapping relationship between the customer type and the risk score; uses the pre-trained first fraud risk identification model to analyze the current transaction information of the target customer to obtain the second risk score; calculates the third risk score of the target customer according to the historical transaction information and the pre-trained second fraud risk identification model; and performs weighted summation on the first risk score, the second risk score and the third risk score to obtain the final risk score of the target customer and feeds it back to the client. In the present invention, by accepting a request to detect the insurance fraud risk of a target customer, the customer information, current transaction information and historical transaction information of the target customer are obtained, the customer type is determined according to the customer information, and the first risk score is determined according to the mapping relationship between the customer type and the risk score; the current transaction information of the target customer is analyzed using the first fraud risk identification model to obtain a second risk score; the historical transaction information is analyzed using the second fraud risk identification model to obtain a third risk score; finally, the three risk scores are weighted and summed to obtain the final risk score and fed back to the client. Through multi-dimensional probability calculation, the limitations of a single dimension are avoided, which helps to reduce false positives and false negatives, thereby improving the accuracy of insurance anti-fraud. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0032] Reference Figure 2 FIG. 1 is a flow chart of an insurance anti-fraud method provided by an embodiment of the present invention. In the embodiment of the present invention, the insurance anti-fraud method includes the following steps S1-S5:

[0033] S1. After receiving a request from a client to detect insurance fraud risk of a target customer, obtain relevant information of the target customer, wherein the relevant information includes customer information, current transaction information, and historical transaction information.

[0034] In one embodiment, in an insurance anti-fraud system, accurately assessing a customer's risk score is key to ensuring the healthy development of the insurance market. When the system receives a request from a client to detect the insurance fraud risk of a target customer, it is necessary to fully obtain and analyze the relevant information of the target customer to improve the accuracy and efficiency of fraud detection. The customer information refers to the personal data of the target customer, including personal information, identity information, and occupational information. The current transaction information refers to the data of the insurance transaction behavior currently proposed by the target customer, including communication records, claims information, and payment records. The historical transaction information refers to the historical insurance transaction data of the target customer, including insurance history records, claims history records, violation history records, and credit history records.

[0035] Specifically, in order to fully obtain the target customer's information data, the insurance anti-fraud system will use a variety of data acquisition methods:

[0036] Internal system: obtain data from the insurance company's internal customer management system, insurance system, claims system, etc.

[0037] External Data Providers: Obtain data from external data providers such as credit rating agencies, public record databases, social media platforms, etc.

[0038] Customer authorization: With the customer's authorization, relevant information is obtained through third-party institutions such as banks and telecom operators.

[0039] Manual investigation: When necessary, obtain detailed customer information through manual investigation and interviews.

[0040] S2. Analyze the customer type of the target customer according to the customer information, and determine a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score.

[0041] In one embodiment, analyzing the customer type of the target customer according to the customer information includes:

[0042] Obtaining customer information of the target customer and determining whether the target customer is a first-time insured customer;

[0043] If the target customer is a first-time insured customer, then the target customer is determined to be a first-time insured customer;

[0044] If the target customer is not insuring for the first time, the average insured amount of the target customer is obtained, and the average insured amount is compared with a preset amount threshold;

[0045] If the average insurance amount of the target customer is greater than or equal to the amount threshold, the target customer is judged to be a high-value customer;

[0046] If the average insurance amount of the target customer is less than the amount threshold, the target customer is judged to be a low-value customer.

[0047] Specifically, the customer's insurance history record is obtained from the customer management system, insurance system, and claims system within the insurance system, as well as external credit rating agencies, public record databases and other data providers. If the target customer has no insurance record at any insurance company, the customer is determined to be a first-time insurance customer. If the target customer has an insurance record, the target customer's insurance amount record is obtained from the insurance system within the insurance system and a third-party agency, and the target customer's average insurance amount is calculated, and the target customer's average insurance amount is compared with a preset amount threshold, and the target customer is marked as a high-value customer or a low-value customer based on the comparison result.

[0048] In one embodiment, the construction of the mapping relationship between the predetermined customer type and the risk score includes:

[0049] Acquire label data of all customers from a preset customer database, wherein the label data includes customer type labels and risk labels, and the risk labels are divided into risky labels and non-risky labels;

[0050] Count the total number of each customer type and the number of customers with risk tags in each customer type;

[0051] Calculate the ratio of the number of customers with risk labels in each customer type to the total number of customers of the corresponding customer type, and obtain the corresponding risk score for each customer type;

[0052] A mapping relationship between the customer type and the risk score is constructed based on the customer type and the risk score corresponding to the customer type.

[0053] In this embodiment, the preset customer database is used to store relevant data of all customers in the insurance system. And the customer database will label customers who have committed fraudulent transactions with risk tags. This embodiment obtains customer type tags and risk tags of all customers from the customer database, such as first-time insurance customer tags, high-value customer tags, and low-value customer tags, and counts the total number of customers corresponding to each customer tag and the number of fraudulent customers in each customer tag, and then calculates the proportion of fraudulent customers in each customer type, and uses the proportion as the risk score corresponding to each customer type, thereby constructing a mapping relationship between the customer type and the risk score.

[0054] Among them, the mapping relationship between the customer type and the risk score can be: [first-time insurance customer—risk score is 20 points], [high-value customer—risk score is 10 points], [low-value customer—risk score is 30 points].

[0055] It is understandable that there is no absolute correlation between customer type and fraud, but the target customers can be initially scored by analyzing the proportion of fraudulent customers in each customer type. Therefore, this embodiment determines the customer type based on customer information, and performs a preliminary screening of target customers based on the mapping relationship between the predetermined customer type and the risk score, which can help insurance companies give priority to customers with higher fraud risks.

[0056] S3. Analyze the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score.

[0057] In one embodiment, the analyzing the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain the second risk score includes:

[0058] Acquire current transaction information of the target customer;

[0059] Extracting features from the current transaction information to obtain first transaction feature data, and converting the first transaction feature data into a first transaction feature vector;

[0060] Calculate the distance between the first transaction feature vector and the feature vector corresponding to the center point of the target cluster representing normal transaction features acquired in advance by using the pre-trained first fraud risk identification model;

[0061] A second risk score of the target customer is calculated according to the distance.

[0062] In this embodiment, the current transaction information refers to the insurance transaction behavior currently proposed by the target customer. It is understandable that if the target customer has completed the insurance transaction behavior, the current transaction information can also be the target customer's most recent insurance transaction behavior. The current transaction information includes customer communication times information (referring to the number of times the customer communicates with the customer service), transaction amount information, transaction page browsing time information (referring to the time it takes the customer to browse the page), transaction page browsing times information, etc. It is understandable that financial fraud behavior usually has a shorter browsing time for the transaction page and fewer browsing times for the transaction page. For example, when a user account is stolen, the hacker usually browses a certain financial insurance with a purpose, and the browsing time and browsing times will be significantly lower than normal transaction behavior. The current transaction feature data includes customer communication times feature, transaction amount feature, transaction page browsing time feature, and transaction page browsing times feature. For example, the target customer's customer communication times are 0 times, the transaction amount feature is 5000, the transaction page browsing time is 30 seconds, and the transaction page browsing times are 1 time. The first fraud model is a clustering model, and the target cluster is a feature that characterizes financial insurance transaction behavior, which is obtained by pre-training the clustering model with large-scale sample data. When the clustering model is trained, similar features will be clustered to form one or more target clusters. For example, the target cluster 1 is the number of customer communications, the target cluster 2 is the transaction amount, the target cluster 3 is the transaction page browsing time, and the target cluster 4 is the transaction page browsing number. Since most of the sample data used are normal transaction behaviors, the cluster center of the formed target cluster is the feature that represents normal transaction behavior. For example, the cluster center of target cluster 1 is the number of customer communications of 2 times, the cluster center of target cluster 2 is the transaction amount of 4000, the cluster center of target cluster 3 is the transaction page browsing time of 120 seconds, and the cluster center of target cluster 4 is the transaction page browsing number of 3 times.

[0063] The pre-trained first fraud risk identification model is a clustering algorithm model, such as the K-means clustering algorithm and the Gaussian mixture model. The distance generally refers to the Euclidean distance, which is not limited here. Taking the K-means clustering algorithm as an example, the K-means clustering algorithm can be used to identify abnormal points or outliers in the data. In financial and insurance transaction data, normal transaction behaviors usually form some relatively dense clusters, while abnormal transaction behaviors (such as fraudulent behaviors) may be far away from these clusters, becoming isolated points or forming small, irregular clusters. Therefore, if the distance between the current transaction feature vector and the target cluster in the first fraud risk identification model is smaller, it means that the current transaction behavior of the target customer is more similar to the normal transaction behavior; if the distance between the current transaction feature vector and the target cluster in the first fraud risk identification model is larger, it means that the current transaction behavior of the target customer is less similar to the normal transaction behavior.

[0064] In one embodiment, the first fraud usage model calculates a second risk score for the target customer using a risk scoring algorithm and the distance, and the risk scoring formula is:

[0065]

[0066] Among them, P(C i |x) represents the risk score of the data point x belonging to the cluster center Ci of the target cluster, x is the normalized current transaction information feature of the target customer, Ci represents the cluster center of the i-th target cluster, K represents the number of cluster centers of the target cluster, j is the index of the cluster center of the target cluster, e represents the base of the natural logarithm, d(x,C i ) represents the distance from the data point x to the cluster center point Ci of the target cluster.

[0067] Specifically, after calculating the risk scores of the first transaction feature vector and the cluster center points of all target clusters, all risk scores are weighted and summed to obtain the second risk score of the target customer. The second risk score can be calculated based on all risk scores by averaging the values, or by setting corresponding weights for each target cluster and summing the values ​​based on the weights. This embodiment does not limit this.

[0068] In this embodiment, the second risk score of the target customer is calculated based on the distance between the current transaction feature vector of the target customer and the cluster centers of multiple target clusters. Generally, the farther the distance between the current transaction feature vector and the cluster center of the target cluster, the higher the risk score of the target customer.

[0069] In one embodiment, the training of the first fraud risk identification model includes:

[0070] Acquire an original data sample set about financial transaction behavior, wherein the original data sample set includes risky transaction data samples and normal transaction data samples;

[0071] Extracting initial sample features from the original data sample set;

[0072] Normalizing the initial sample features and using them as training data;

[0073] The training data is input into a pre-built clustering algorithm model to perform clustering training on the clustering algorithm model to obtain a target cluster representing normal transaction characteristics after clustering and the first fraud risk identification model.

[0074] In this embodiment, the financial transaction behavior refers to the transaction behavior related to financial insurance, and the original data sample set includes risky transaction data samples and normal transaction data samples related to financial insurance behavior. Among them, the financial transaction data can come from internal systems and external data providers. Then, the initial sample features of the risky transaction data samples and normal transaction data samples can be extracted from the aforementioned original data sample set, such as the number of customer communication features, transaction amount features, transaction page browsing time features, and transaction page browsing times features. Then, the aforementioned initial sample features can be normalized using a predetermined function (such as a Z-score normalization function, a Sigmoid function, a (0, 1) normalization function, etc.) and used as training data.

[0075] In this embodiment, the clustering model can quickly classify the current transaction information into different clusters and evaluate the abnormality of the transaction in real time. In addition, the clustering model is an unsupervised learning method that does not require pre-labeled label data and is suitable for real-time transaction monitoring. It can identify abnormal fraudulent behavior without clear risk labels. Therefore, this embodiment uses the clustering model to identify the current transaction information of the target customer, thereby quickly identifying abnormal transactions, thereby improving the efficiency and accuracy of identifying financial fraud.

[0076] In one embodiment, the clustering training of the clustering algorithm model includes:

[0077] The training data is clustered using the K-means clustering algorithm, and the following operations are performed cyclically during the clustering process:

[0078] Calculating the inter-cluster distances of each trained cluster and the intra-cluster distances between elements in each cluster to obtain clustering results; and

[0079] The number of trained clusters is adjusted according to the clustering result until the target cluster is output.

[0080] In this embodiment, in some implementation scenarios, the Dunn index function can be used to characterize the inter-cluster distance of each cluster, and the compactness index function can be used to characterize the intra-cluster distance. It can be understood that the K-means clustering algorithm, the Dunn index and the compactness index are used as examples for exemplary description, and the technical scheme of the present invention is not limited. Among them, the target cluster represents normal financial transaction behavior and the inter-cluster distance of each cluster in the target cluster and the intra-cluster distance between the elements in each cluster meet the predetermined evaluation index to complete the clustering training of the clustering algorithm model. In some embodiments, as described above, the Dunn index and the compactness index can be used to characterize the inter-cluster distance and the intra-cluster distance respectively. The aforementioned predetermined evaluation index can be represented by (2DVI×CP / (DVI+CP))max, where DVI represents the Dunn index, CP represents the compactness index, and the maximum value of the formula can be used as the aforementioned predetermined evaluation index. It can be understood that the description of the predetermined evaluation index here is only an exemplary description, and the technical scheme of the present invention is not limited thereto.

[0081] In one embodiment, the present invention may also set a scoring warning mechanism. When the second risk score of the target customer reaches a preset warning threshold (for example, the warning threshold is 95 points), a warning prompt is issued to the client. The client takes corresponding warning measures when receiving the warning prompt, for example, by limiting the network speed or popping up a verification window multiple times to delay the target customer's transaction behavior, or directly suspending the target customer's transaction rights. In addition, after the warning prompt is issued to the client, subsequent risk assessments can be continued to prevent misjudgment. It is understandable that if the current transaction information of the target customer is the most recent transaction information that the target customer has completed, the subsequent risk assessment will continue until the final risk score is obtained. In this embodiment, the real-time analysis of the current transaction information can capture the immediate abnormal behavior and ensure that a response is made at the first time the transaction occurs. This is very important for preventing immediate fraud, especially in online transactions or automated processing scenarios.

[0082] S4. Calculate a third risk score for the target customer based on the historical transaction information and the pre-trained second fraud risk identification model.

[0083] Since only analyzing the current transaction information of the target customer may result in misjudgment or inaccurate risk scoring, the present invention further analyzes the historical transaction information of the target customer through a second fraud risk identification model.

[0084] In one embodiment, the calculating the third risk score of the target customer based on the historical transaction information and the pre-trained second fraud risk identification model includes:

[0085] Obtaining historical transaction information of the target customer;

[0086] Extracting second transaction feature data of the target customer from the historical transaction information, and converting the second transaction feature data into a second transaction feature vector;

[0087] The second transaction feature vector is input into a pre-trained second fraud risk identification model, and the third risk score of the target customer is obtained.

[0088] In this embodiment, the historical transaction information of the target customer includes the customer's transaction amount record, login record, transaction channel record and account activity record. The transaction amount record refers to the historical transaction order data of the target customer obtained from the insurance system, including the historical insurance amount and the historical claim amount, etc.; the login record refers to the login time, login location and login device corresponding to each transaction order of the target customer; the transaction channel record refers to the transaction method used by the target customer for each transaction order, such as cash, ATM or online payment, etc.; the account activity record refers to the fund change record in the target customer's account, such as large-scale fund transfer and frequent fund inflow and outflow, etc. The pre-trained second fraud risk identification model can be a neural network model, a random forest model or a support vector machine model, etc. In this embodiment, the neural network model is taken as an example. The neural network model has a strong nonlinear modeling ability, can capture complex data patterns and hidden relationships, and is suitable for processing large-scale historical transaction data. The neural network model is trained on large-scale data. As the amount of data increases, the generalization ability and accuracy of the model will continue to improve, and it can more accurately identify fraudulent behavior, thereby improving the accuracy of identifying financial fraud behavior.

[0089] In one embodiment, the training of the second fraud risk identification model includes:

[0090] A. Obtain a first sample set and a second sample set about financial transaction behaviors, wherein the first sample set is samples without risk labels, and the second sample set is samples with risk labels;

[0091] Specifically, the first sample set is sample users without labeled labels, and the second sample set is sample users with labeled labels. Exemplarily, the number of the first sample set without labeled labels and the number of the second sample set with labeled labels can be the same or different. In general scenarios, the number of the first sample set without labels is greater than the number of the second sample set with labels. Since the historical transaction information is not standardized, it is not conducive to computer automated processing. Therefore, the vectorization of data is used to convert the non-standard data into a vector form with a consistent format that is convenient for computer processing.

[0092] B. performing feature extraction and vector conversion on the first sample set to obtain a first feature vector set, and performing feature extraction and vector conversion on the second sample set to obtain a second feature vector set;

[0093] The method of extracting the feature vector includes:

[0094] For each first sample set, a feature value of the first sample set under a plurality of preset operation behavior features is determined according to historical transaction information of the first sample set.

[0095] According to the feature values ​​of the first sample set under a plurality of preset operation behavior features, a first feature vector that can be used to characterize the operation behavior features of the first sample set is constructed.

[0096] For each second sample set, a feature value of the second sample set under a plurality of preset operation behavior features is determined according to historical transaction information of the second sample set.

[0097] According to the feature values ​​of the second sample set under a plurality of preset operation behavior features, a second feature vector that can be used to characterize the operation behavior features of the second sample set is constructed.

[0098] Specifically, the historical transaction information of the sample set includes multiple preset operation behavior features, such as fund transfer, remote login, frequent balance changes, etc. Among them, the corresponding numerical representation is directly used for the numerical feature, and the one-hot encoding method is used for the category feature, that is, each category feature corresponds to a vector composed of 0 and 1, and the number of categories corresponds to the dimension of the vector, that is, one category corresponds to one dimension of the vector. When the preset operation behavior feature is a certain category, the vector position corresponding to the category is 1, and the other parts are all set to 0. For example, the preset operation behavior feature "Is it the device information commonly used by the user" includes two categories, namely "yes" and "no", then the preset operation behavior feature "Is it the device information commonly used by the user" uses a two-bit hot-single encoding method, assuming that "yes" is "10" and "no" is "01".

[0099] C. Preliminarily training the pre-constructed neural network model using the second feature vector set to obtain an initial neural network model;

[0100] D. inputting the first feature vector set into the initial neural network model to obtain a risk prediction result;

[0101] Specifically, the risk prediction result refers to the risk score of the initial neural network model for the first feature vector set. After obtaining the initial neural network model trained by the labeled sample, the first feature vector set is then input into the initial neural network model. The model will output the risk prediction score (0-10 points) for each sample. Based on these scores, a pseudo-label can be generated for each unlabeled sample. For example, if the risk prediction score of the sample is greater than 8, it can be marked as "risk"; if it is less than 2, it is marked as "no risk". In order to ensure the quality of the pseudo-label, the present embodiment can also set a scoring threshold, and only select samples that meet the scoring threshold requirements and include them in subsequent training, which helps reduce the risk of mislabeling, thereby improving the accuracy of the model. For example, the scoring threshold is set to 9 and 1, and samples with a risk score greater than or equal to 9 or samples with a risk score less than or equal to 1 are selected.

[0102] E. Based on the first feature vector set and the corresponding risk prediction result, and the second feature vector set containing the risk label, the initial neural network model is iteratively trained once to obtain an intermediate neural network model;

[0103] In each iteration, the model is trained not only with the second sample set with annotated labels, but also with the first sample set that has been screened. Through multiple iterations of training, the fraud prediction results obtained by using the trained neural network model to predict the fraud results of the historical transaction information of the first sample set will be closer and closer to the real labels, thereby making the model training more and more accurate, and finally obtaining a fraud identification model that meets the accuracy requirements.

[0104] F. Using a pre-acquired test set to verify the intermediate neural network model to obtain a loss value, and comparing the loss value with a preset loss threshold to obtain a comparison result;

[0105] In one embodiment, the step of using a pre-acquired test set to verify the intermediate neural network model to obtain a loss value includes:

[0106] The test set is input into the intermediate neural network model to obtain the predicted label of the test set, and the loss value of the test label and the true label is calculated using a preset loss function.

[0107] Specifically, the test set is a sample set with real labels and does not contain samples repeated in the second sample set. In one embodiment, when obtaining a sample set about financial transaction behavior in advance, the samples with labels can be divided into the second sample set and the test set. The loss value is calculated using a preset loss function, which is used to measure the difference between the model prediction value and the real label. Common loss functions include cross-entropy loss (Cross-Entropy Loss), mean squared error (MSE), etc., which are not limited here.

[0108] G. Determine, based on the comparison result, whether to perform the next round of iterative training on the intermediate neural network model or output the intermediate neural network model as the second fraud identification model.

[0109] Specifically, if the comparison result is that the loss value is greater than the loss threshold, the intermediate neural network model is subjected to the next round of iterative training; if the comparison result is that the loss value is less than or equal to the loss threshold, the intermediate neural network model is output as the second fraud identification model. In one embodiment, the number of iterative training is also preset, and when the iterative training of the model reaches the preset number of times, the intermediate neural network model is directly output as the second fraud identification model.

[0110] In this embodiment, by analyzing historical transaction information, the customer's long-term behavior pattern can be identified, helping insurance companies better understand the customer's risk characteristics. The accumulation of historical data enables the model to be continuously optimized and improves the accuracy of fraud detection.

[0111] S5. Perform a weighted summation on the first risk score, the second risk score and the third risk score to obtain a final risk score of the target customer and feed it back to the client.

[0112] In one embodiment, the present application sets weights for three risk scores, then multiplies the three probabilities by their respective weights and sums the three multiplied results. For example, the weights corresponding to the first risk score, the second risk score, and the third risk score are 0.2, 0.4, and 0.4, respectively. It is understandable that the final insurance probability can also be calculated by adding the three risk scores and averaging them according to actual needs, which is not limited here.

[0113] In this embodiment, the final risk score is obtained by weighted summing the first, second and third risk scores. This comprehensive evaluation method considers information from multiple dimensions, avoids the limitations of a single dimension, and thus improves the accuracy of fraud detection.

[0114] In the above technical solution provided by this embodiment, by accepting a request to detect the insurance fraud risk of a target customer, the customer information, current transaction information and historical transaction information of the target customer are obtained, the customer type is determined according to the customer information, and the first risk score is determined according to the mapping relationship between the customer type and the risk score; the current transaction information of the target customer is analyzed using the first fraud risk identification model to obtain the second risk score; the historical transaction information is analyzed using the second fraud risk identification model to obtain the third risk score; finally, the three risk scores are weighted and summed to obtain the final risk score and fed back to the client. Through multi-dimensional probability calculation, the limitations of a single dimension are avoided, which helps to reduce false positives and false negatives, thereby improving the accuracy of insurance anti-fraud.

[0115] like Figure 3 FIG. 1 is a schematic diagram of the structure of an insurance anti-fraud device provided by an embodiment of the present invention.

[0116] The insurance anti-fraud device 100 of the present invention can be installed in an electronic device. According to the functions to be implemented, the insurance anti-fraud device can include a relevant information acquisition module 101, a first risk scoring module 102, a second risk scoring module 103, a third risk scoring module 104, and a risk scoring return module 105. The module of the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, and is stored in the memory of the electronic device.

[0117] In this embodiment, the functions of each module / unit are as follows:

[0118] Relevant information acquisition module 101: after receiving a request from a client to detect insurance fraud risk of a target customer, acquire relevant information of the target customer, the relevant information including customer information, current transaction information and historical transaction information;

[0119] The first risk scoring module 102 is used to analyze the customer type of the target customer according to the customer information, and determine the first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score;

[0120] The second risk scoring module 103 is used to analyze the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score;

[0121] The third risk scoring module 104 is used to calculate the third risk score of the target customer according to the historical transaction information and the pre-trained second fraud risk identification model;

[0122] The risk score returning module 105 is used to perform a weighted summation on the first risk score, the second risk score and the third risk score to obtain a final risk score of the target customer and feed it back to the client.

[0123] In one embodiment, when performing the step of analyzing the customer type of the target customer according to the customer information, the first risk scoring module 102 is configured to:

[0124] Obtaining customer information of the target customer and determining whether the target customer is a first-time insured customer;

[0125] If the target customer is a first-time insured customer, then the target customer is determined to be a first-time insured customer;

[0126] If the target customer is not insuring for the first time, the average insured amount of the target customer is obtained, and the average insured amount is compared with a preset amount threshold;

[0127] If the average insurance amount of the target customer is greater than or equal to the amount threshold, the target customer is judged to be a high-value customer;

[0128] If the average insurance amount of the target customer is less than the amount threshold, the target customer is judged to be a low-value customer.

[0129] In one embodiment, when executing the construction of the mapping relationship between the predetermined customer type and the risk score, the first risk scoring module 102 is used to:

[0130] Acquire label data of all customers from a preset customer database, wherein the label data includes customer type labels and risk labels, and the risk labels are divided into risky labels and non-risky labels;

[0131] Count the total number of each customer type and the number of customers with risk tags in each customer type;

[0132] Calculate the ratio of the number of customers with risk labels in each customer type to the total number of customers of the corresponding customer type, and obtain the corresponding risk score for each customer type;

[0133] A mapping relationship between the customer type and the risk score is constructed based on the customer type and the risk score corresponding to the customer type.

[0134] In one embodiment, when the second risk scoring module 103 performs the analysis of the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain the second risk score, it is used to:

[0135] Acquire current transaction information of the target customer;

[0136] Extracting features from the current transaction information to obtain first transaction feature data, and converting the first transaction feature data into a first transaction feature vector;

[0137] Calculate the distance between the first transaction feature vector and the feature vector corresponding to the center point of the target cluster representing normal transaction features acquired in advance by using the pre-trained first fraud risk identification model;

[0138] A second risk score of the target customer is calculated according to the distance.

[0139] In one embodiment, when executing the training of the first fraud risk identification model, the third risk scoring module 103 is used to:

[0140] Acquire an original data sample set about financial transaction behavior, wherein the original data sample set includes risky transaction data samples and normal transaction data samples;

[0141] Extracting initial sample features from the original data sample set;

[0142] Normalizing the initial sample features and using them as training data;

[0143] The training data is input into the pre-constructed clustering algorithm model to perform clustering training on the clustering algorithm model, and obtain the clustered target class cluster that represents normal transaction characteristics and the first fraud risk identification model.

[0144] In one embodiment, the third risk scoring module 104 is used to: when performing the calculation of the third risk score of the target customer based on the historical transaction information and the pre-trained second fraud risk identification model:

[0145] Obtaining historical transaction information of the target customer;

[0146] Extracting second transaction feature data of the target customer from the historical transaction information, and converting the second transaction feature data into a second transaction feature vector;

[0147] The second transaction feature vector is input into a pre-trained second fraud risk identification model, and the third risk score of the target customer is obtained.

[0148] In one embodiment, when executing the training of the second fraud risk identification model, the third risk scoring module 104 is used to:

[0149] Acquire a first sample set and a second sample set about financial transaction behaviors, wherein the first sample set is samples without risk labels, and the second sample set is samples with risk labels;

[0150] Performing feature extraction and vector conversion on the first sample set to obtain a first feature vector set, and performing feature extraction and vector conversion on the second sample set to obtain a second feature vector set;

[0151] Performing preliminary training on the pre-constructed neural network model using the second feature vector set to obtain an initial neural network model;

[0152] Inputting the first feature vector set into the initial neural network model to obtain a risk prediction result;

[0153] Based on the first feature vector set and the corresponding risk prediction result, and the second feature vector set containing the annotated risk labels, the initial neural network model is iteratively trained once to obtain an intermediate neural network model;

[0154] Using a pre-acquired test set to verify the intermediate neural network model to obtain a loss value, and comparing the loss value with a preset loss threshold to obtain a comparison result;

[0155] According to the comparison result, determine whether to perform the next round of iterative training on the model or output the intermediate neural network model as the second fraud identification model.

[0156] The present invention provides an insurance anti-fraud device, which receives a request for detecting the insurance fraud risk of a target customer, obtains the customer information, current transaction information and historical transaction information of the target customer, determines the customer type according to the customer information, and determines the first risk score according to the mapping relationship between the customer type and the risk score; uses the first fraud risk identification model to analyze the current transaction information of the target customer to obtain the second risk score; uses the second fraud risk identification model to analyze the historical transaction information to obtain the third risk score; finally, the three risk scores are weighted and summed to obtain the final risk score and fed back to the client. Through multi-dimensional probability calculation, the limitations of a single dimension are avoided, which helps to reduce false positives and false negatives, thereby improving the accuracy of insurance anti-fraud.

[0157] The specific definition of the insurance anti-fraud device can be found in the definition of the insurance anti-fraud method above, which will not be repeated here. Each module in the above-mentioned insurance anti-fraud device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0158] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a service-side of an insurance anti-fraud method.

[0159] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a function or step on the insurance anti-fraud client side

[0160] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0161] After receiving a request from a client to detect insurance fraud risk of a target customer, obtaining relevant information of the target customer, the relevant information including customer information, current transaction information and historical transaction information;

[0162] Analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score;

[0163] Analyzing the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score;

[0164] Calculate a third risk score for the target customer based on the historical transaction information and the pre-trained second fraud risk identification model;

[0165] The first risk score, the second risk score and the third risk score are weightedly summed to obtain a final risk score of the target customer and feed it back to the client.

[0166] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0167] After receiving a request from a client to detect insurance fraud risk of a target customer, obtaining relevant information of the target customer, the relevant information including customer information, current transaction information and historical transaction information;

[0168] Analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score;

[0169] Analyzing the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score;

[0170] Calculate a third risk score for the target customer based on the historical transaction information and the pre-trained second fraud risk identification model;

[0171] The first risk score, the second risk score and the third risk score are weightedly summed to obtain a final risk score of the target customer and feed it back to the client.

[0172] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0173] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0174] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0175] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should be included in the protection scope of the present invention. It should be noted that if software tools or components other than those of the company appear in the embodiments of this application, they are only used for example introduction and do not represent actual use.

Claims

1. An insurance anti-fraud method, characterized in that: The method comprises: After receiving a request from a client to detect insurance fraud risk of a target customer, obtaining relevant information of the target customer, the relevant information including customer information, current transaction information and historical transaction information; Analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score; Analyzing the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score; Calculate a third risk score for the target customer based on the historical transaction information and the pre-trained second fraud risk identification model; The first risk score, the second risk score and the third risk score are weightedly summed to obtain a final risk score of the target customer and feed it back to the client.

2. The insurance anti-fraud method according to claim 1, characterized in that: Analyzing the customer type of the target customer according to the customer information includes: Determining whether the target customer is a first-time insured customer based on the customer information of the target customer; If the target customer is a first-time insured customer, determining the customer type of the target customer as a first-time insured customer; If the target customer is not insuring for the first time, the average insured amount of the target customer is obtained, and the average insured amount is compared with a preset amount threshold; If the average insurance amount of the target customer is greater than or equal to the amount threshold, the customer type of the target customer is determined to be a high-value customer; If the average insured amount of the target customer is less than the amount threshold, the customer type of the target customer is determined to be a low-value customer.

3. The insurance anti-fraud method according to claim 1, characterized in that: The construction of the mapping relationship between the predetermined customer type and the risk score includes: Acquire label data of all customers from a preset customer database, wherein the label data includes customer type labels and risk labels, and the risk labels are divided into risky labels and non-risky labels; Count the total number of each customer type and the number of customers with risk tags in each customer type; Calculate the ratio of the number of customers with risk labels in each customer type to the total number of customers of the corresponding customer type, and obtain the corresponding risk score for each customer type; A mapping relationship between the customer type and the risk score is constructed based on the customer type and the risk score corresponding to the customer type.

4. The insurance anti-fraud method according to claim 1, characterized in that: The using the pre-trained first fraud risk identification model to analyze the current transaction information of the target customer to obtain a second risk score includes: Obtaining current transaction information of the target customer; Extracting features from the current transaction information to obtain first transaction feature data, and converting the first transaction feature data into a first transaction feature vector; Calculate the distance between the first transaction feature vector and the feature vector corresponding to the center point of the target cluster representing normal transaction features acquired in advance by using the pre-trained first fraud risk identification model; A second risk score of the target customer is calculated according to the distance.

5. The insurance anti-fraud method according to claim 1, characterized in that: The training of the first fraud risk identification model includes: Acquire an original data sample set about financial transaction behavior, wherein the original data sample set includes risky transaction data samples and normal transaction data samples; Extracting initial sample features from the original data sample set; Normalizing the initial sample features and using them as training data; The training data is input into a pre-built clustering algorithm model to perform clustering training on the clustering algorithm model, thereby obtaining a clustered target cluster representing normal transaction characteristics and the first fraud risk identification model.

6. The insurance anti-fraud method according to claim 1, characterized in that: The step of calculating a third risk score of the target customer based on the historical transaction information and the pre-trained second fraud risk identification model includes: Obtaining historical transaction information of the target customer; Extracting second transaction feature data of the target customer from the historical transaction information, and converting the second transaction feature data into a second transaction feature vector; The second transaction feature vector is input into a pre-trained second fraud risk identification model, and the third risk score of the target customer is obtained.

7. The insurance anti-fraud method according to claim 1, characterized in that: The training of the second fraud risk identification model includes: Acquire a first sample set and a second sample set about financial transaction behaviors, wherein the first sample set is samples without risk labels, and the second sample set is samples with risk labels; Performing feature extraction and vector conversion on the first sample set to obtain a first feature vector set, and performing feature extraction and vector conversion on the second sample set to obtain a second feature vector set; Performing preliminary training on the pre-constructed neural network model using the second feature vector set to obtain an initial neural network model; Inputting the first feature vector set into the initial neural network model to obtain a risk prediction result; Based on the first feature vector set and the corresponding risk prediction result, and the second feature vector set containing the annotated risk labels, the initial neural network model is iteratively trained once to obtain an intermediate neural network model; Using a pre-acquired test set to verify the intermediate neural network model to obtain a loss value, and comparing the loss value with a preset loss threshold to obtain a comparison result; According to the comparison result, determine whether to perform the next round of iterative training on the model or output the intermediate neural network model as the second fraud identification model.

8. An insurance anti-fraud device, characterized in that: The device comprises: Relevant information acquisition module: used to acquire relevant information of a target customer after receiving a request sent by a client to detect the insurance fraud risk of the target customer, wherein the relevant information includes customer information, current transaction information and historical transaction information; A first risk scoring module: used for analyzing the customer type of the target customer according to the customer information, and determining a first risk score of the target customer according to a predetermined mapping relationship between the customer type and the risk score; A second risk scoring module: used to analyze the current transaction information of the target customer using the pre-trained first fraud risk identification model to obtain a second risk score; A third risk scoring module: used for calculating a third risk score of the target customer based on the historical transaction information and the pre-trained second fraud risk identification model; Risk score return module: used to perform weighted summation on the first risk score, the second risk score and the third risk score to obtain a final risk score of the target customer and feed it back to the client.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the insurance anti-fraud method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the insurance anti-fraud method as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Fraudulent transaction judgement method and device, computer device and storage medium

    CN108717638A

  • A claim settlement anti-fraud risk control method and device

    CN109191312A

  • Risk customer identification method and device and server

    CN110675268A

  • Financial customer fraud risk identification method and device

    CN111275546A

  • Anti-fraud method, device and equipment

    CN113744054A