A fraud risk identification method and device based on graph neural network

The graph neural network method addresses the issue of isolated data islands in vehicle insurance fraud detection by constructing association matrices and predicting fraud probabilities, enhancing the accuracy of fraud risk identification.

CN115953172BActive Publication Date: 2025-07-15ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211625194.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-07-15
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

The existing fraud risk identification methods in the field of auto insurance claims fail to effectively correlate the historical data in the database, resulting in inaccurate identification results.

Method used

Through a graph neural network-based method, the association relationship adjacency matrix and feature matrix of the events to be identified are obtained, and the pre-trained graph neural network is used to calculate the fraud probability, thereby identifying the fraud risk.

Benefits of technology

The accuracy of fraud risk identification is improved and the inaccurate identification results caused by the failure to effectively associate database historical data in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115953172B_ABST
    Figure CN115953172B_ABST
Patent Text Reader

Abstract

The present application relates to a fraud risk identification method and device based on a graph neural network. The method includes: obtaining a first association relationship adjacency matrix of an event to be identified according to a data table; the data table includes data of the event to be identified, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be identified and other events in the data table; obtaining a first feature matrix of the event to be identified according to the data table and the first association relationship adjacency matrix; the first feature matrix is used to identify the data of the events in the data table; obtaining the fraud probability of the event to be identified according to the first association relationship adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of an event; determining the fraud risk of the event to be identified according to the fraud probability. Through the present application, the problem that the existing fraud risk identification method in the field of auto insurance claims does not associate historical data in the database, resulting in inaccurate identification results, is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a fraud risk identification method and device based on a graph neural network. Background Art

[0002] With the rapid development of computer technology, a large amount of business data is generated in actual business and stored in a computer database. There is a certain correlation between the actual data generated by users with fraudulent behaviors in the processes of motor vehicle insurance reporting, accident occurrence, and claim settlement. It is becoming increasingly important to timely discover potential fraud risk points from the massive business data for determining whether a specific case is fraudulent. Business institutions store business data in the database according to their respective categories. However, since different data tables are stored independently of each other, it is difficult to connect them one by one through primary keys, forming one data island after another, and it is difficult to simply and batch compare related cases.

[0003] Existing fraud risk identification methods in the field of motor vehicle insurance claim settlement do not associate historical data in the database, resulting in inaccurate identification results.

[0004] Regarding the problem of inaccurate fraud risk identification results in the existing technology, no effective solution has been proposed yet. Summary of the Invention

[0005] In this embodiment, a fraud risk identification method and device based on a graph neural network are provided to solve the problem of inaccurate fraud risk identification results in the existing technology.

[0006] In the first aspect, in this embodiment, a fraud risk identification method based on a graph neural network is provided. The method includes:

[0007] Obtain a first association relationship adjacency matrix of the event to be identified according to the data table; the data table includes data of the event to be identified, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be identified and other events in the data table;

[0008] Obtain a first feature matrix of the event to be identified according to the data table and the first association relationship adjacency matrix; the first feature matrix is used to identify the data of the events in the data table;

[0009] Obtain the fraud probability of the event to be identified according to the first association relationship adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of an event;

[0010] Determine the fraud risk of the event to be identified according to the fraud probability.

[0011] In some of these embodiments, obtaining the first associated relationship adjacency matrix of the event to be recognized according to the data table includes:

[0012] Obtaining the associated relationship graph of the event to be recognized according to the data table; the associated relationship graph is used to identify the associated relationship between the event to be recognized and other events in the data table;

[0013] Generating the first associated relationship adjacency matrix of the event to be recognized according to the associated relationship graph.

[0014] In some of these embodiments, obtaining the first feature matrix of the event to be recognized according to the data table and the first associated relationship adjacency matrix includes:

[0015] Adjusting the record data of the data table according to the event order in the first associated relationship adjacency matrix;

[0016] Obtaining the first feature matrix of the event to be recognized according to the adjusted data table.

[0017] In some of these embodiments, obtaining the first feature matrix of the event to be recognized according to the adjusted data table includes:

[0018] Obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table, obtaining the corresponding eigenvalue according to the feature vector, and generating the first feature matrix according to the eigenvalue.

[0019] In some of these embodiments, obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table includes:

[0020] Obtaining the feature vector corresponding to the event to be recognized from the adjusted data table according to the associated relationship between the event to be recognized and other events.

[0021] In some of these embodiments, the method further includes:

[0022] Obtaining the second associated relationship adjacency matrix of the training event according to the data table;

[0023] Obtaining the second feature matrix of the training event according to the data table and the second associated relationship adjacency matrix;

[0024] Training the graph neural network by using the second associated relationship adjacency matrix and the second feature matrix.

[0025] In some of these embodiments, training the graph neural network by using the second associated relationship adjacency matrix and the second feature matrix includes:

[0026] Input the second association relationship adjacency matrix of the training event and the second feature matrix of the training event into the graph neural network to obtain a predicted value;

[0027] Construct a loss function of the graph neural network according to the predicted value and the label value of the training event;

[0028] Adjust the parameters of the graph neural network according to the loss function.

[0029] In a second aspect, a fraud risk identification device based on a graph neural network is provided in this embodiment. The device includes:

[0030] A first acquisition module, configured to acquire a first association relationship adjacency matrix of an event to be identified according to a data table; the data table includes data of the event to be identified, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be identified and other events in the data table;

[0031] A second acquisition module, configured to acquire a first feature matrix of the event to be identified according to the data table and the first association relationship adjacency matrix; the first feature matrix is used to identify data of events in the data table;

[0032] A third acquisition module, configured to acquire a fraud probability of the event to be identified according to the first association relationship adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to acquire the fraud probability of an event;

[0033] A determination module, configured to determine the fraud risk of the event to be identified according to the fraud probability.

[0034] In a third aspect, an electronic device is provided in this embodiment, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the fraud risk identification method based on a graph neural network described in the first aspect.

[0035] In a fourth aspect, a computer-readable storage medium is provided in this embodiment, on which a computer program is stored. When the computer program is executed by a processor, the steps of the fraud risk identification method based on a graph neural network described in the first aspect are implemented.

[0036] Compared with the prior art, a fraud risk identification method and device based on a graph neural network provided in this embodiment solve the problem that the existing fraud risk identification method in the field of auto insurance claims does not associate historical data in the database, resulting in inaccurate identification results. This is achieved by obtaining a first association relationship adjacency matrix based on historical data in a data table, obtaining a first feature matrix of an event to be identified based on the first association relationship adjacency matrix, obtaining the fraud probability of the event to be identified through a graph neural network based on the first association relationship adjacency matrix and the first feature matrix, and performing risk identification based on the fraud probability.

[0037] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0039] Figure 1 is a hardware structure block diagram of a terminal for executing a fraud risk identification method based on a graph neural network according to an embodiment of the present application;

[0040] Figure 2 is a flowchart of a fraud risk identification method based on a graph neural network according to an embodiment of the present application;

[0041] Figure 3 is a flowchart of obtaining a first association relationship adjacency matrix of an event to be identified according to a data table in an embodiment of the present application;

[0042] Figure 4 is a flowchart of a fraud risk identification method based on a graph neural network in this specific embodiment;

[0043] Figure 5 is a structure block diagram of a fraud risk identification device based on a graph neural network according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To more clearly understand the purpose, technical solution, and advantages of the present application, the present application is described and explained below with reference to the drawings and embodiments.

[0045] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The words such as "connect", "be connected", "couple" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific sorting of the objects.

[0046] The method embodiments provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, when running on a terminal, Figure 1 is a hardware structure block diagram of a terminal that executes a fraud risk identification method based on a graph neural network according to an embodiment of this application. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in Figure 1 ) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 , or have a different configuration from that shown in Figure 1 .

[0047] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to a fraud risk identification method based on a graph neural network in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0048] The transmission device 106 is used to receive or send data via a network. The above network includes the wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0049] In this embodiment, a fraud risk identification method based on a graph neural network is provided. Figure 2 It is a flowchart of a fraud risk identification method based on a graph neural network according to an embodiment of the present application, as Figure 2 shown, and the process includes the following steps:

[0050] Step S210, obtaining a first associated relationship adjacency matrix of the event to be identified according to a data table; the data table includes data of the event to be identified, and the first associated relationship adjacency matrix is used to identify the associated relationship between the event to be identified and other events in the data table.

[0051] Specifically, according to the event identifier of the event to be identified, a data table related to the event to be identified is obtained, and the data table includes data of the event to be identified. A first associated relationship adjacency matrix of the event to be identified is obtained according to the data table, and the first associated relationship adjacency matrix is used to identify the associated relationship between the event to be identified and other events in the data table.

[0052] Exemplarily, the data table is stored in a historical event structured database, and the historical event structured database is used to store information related to historical event auto insurance claims events. The first associated relationship adjacency matrix is an N*N-dimensional matrix, where N is the number of events in the data table, and the element A in the first associated relationship adjacency matrix ij, indicating the correlation between event i and time j. For example, element A ij being 0 indicates that there is no correlation between event i and event j, and element A ij being 1 indicates that there is a correlation between event i and event j.

[0053] Step S220: Obtain the first feature matrix of the event to be recognized according to the data table and the first correlation adjacency matrix; the first feature matrix is used to identify the data of the events in the data table.

[0054] Specifically, according to the first correlation adjacency matrix obtained in step S210, obtain the first feature matrix of the event to be recognized from the data table, and this first feature matrix is used to identify the data of the events in the data table.

[0055] Exemplarily, if the data table records the data of M dimensions of an event, then extract the data of M dimensions of the corresponding event from the data table according to the first correlation adjacency matrix to form this first feature matrix, and the first feature matrix is an N*M-dimensional matrix.

[0056] Step S230: Obtain the fraud probability of the event to be recognized according to the first correlation adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of the event.

[0057] Specifically, input the first correlation adjacency matrix and the first feature matrix into the pre-trained graph neural network to obtain the fraud probability of the event to be recognized. The output of this graph neural network is a probability value, and this probability value is used to identify the probability of fraud existence of the event to be recognized.

[0058] Step S240: Determine the fraud risk of the event to be recognized according to the fraud probability.

[0059] Specifically, determine the fraud risk of the event to be recognized according to the fraud probability output by the graph neural network in step S230. Exemplarily, when the fraud probability is higher, the fraud risk of the event to be recognized is greater; when the fraud probability is lower, the fraud risk of the event to be recognized is smaller.

[0060] In this embodiment, by obtaining the first correlation adjacency matrix according to the historical data in the data table, obtaining the first feature matrix of the event to be recognized according to this first correlation adjacency matrix, obtaining the fraud probability of the event to be recognized through the graph neural network based on the first correlation adjacency matrix and the first feature matrix, and performing risk identification according to this fraud probability, the problem that the existing fraud risk identification method in the field of vehicle insurance claims does not correlate the historical data in the database, resulting in inaccurate identification results, is solved.

[0061] In some of these embodiments, in step S210, a first association relationship adjacency matrix of the event to be recognized is obtained according to the data table, as Figure 3 shown, which includes the following steps:

[0062] Step S211, obtaining an association relationship graph of the event to be recognized according to the data table; the association relationship graph is used to identify the association relationship between the event to be recognized and other events in the data table.

[0063] Specifically, an association relationship graph of the event to be recognized is obtained from the data table of the historical event structured database, and the association relationship graph is used to identify the association relationship between the event to be recognized and other events in the data table. Exemplarily, the nodes of the association relationship graph are used to identify events. If there is an association relationship edge between two nodes, it means that there is an association relationship between the two events.

[0064] Step S212, generating a first association relationship adjacency matrix of the event to be recognized according to the association relationship graph.

[0065] Exemplarily, the first association relationship adjacency matrix is an N*N-dimensional matrix, where N is the number of events in the data table. The element A ij in the first association relationship adjacency matrix represents the association relationship between event i and time j. For example, when the element A ij is 0, it means that there is no association relationship edge between event i and event j, that is, there is no association relationship between event i and event j. When the element A ij is 1, it means that there is an association relationship edge between event i and event j, that is, there is an association relationship between event i and event j.

[0066] In this embodiment, a first association relationship adjacency matrix is generated according to the association relationship graph of the event to be recognized, so as to associate the historical data in the database and improve the accuracy of the recognition result.

[0067] In some of these embodiments, according to the data table and the first association relationship adjacency matrix, a first feature matrix of the event to be recognized is obtained, including: adjusting the record data of the data table according to the event order in the first association relationship adjacency matrix; obtaining the first feature matrix of the event to be recognized according to the adjusted data table.

[0068] Specifically, when generating the first association relationship adjacency matrix according to the association relationship graph, the order of events in the first association relationship adjacency matrix will be inconsistent with the order of events in the data table. For example, when generating the first association relationship adjacency matrix, the order of events with a larger association relationship degree will be adjusted to the front. Exemplarily, the order of events in the data table is N1, N2, N3, N4, N5, and the events represented by the rows and columns of the generated first association relationship adjacency matrix are (N2, N3, N1, N4, N5) and (N2, N3, N1, N4, N5). According to the order of events in the first association relationship adjacency matrix, the record data of the data table is adjusted, and the order of events in the adjusted data table is N2, N3, N1, N4, N5. According to the adjusted data table, the first feature matrix of the event to be recognized is obtained.

[0069] In some of these embodiments, obtaining the first feature matrix of the event to be recognized according to the adjusted data table includes: obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table, obtaining the corresponding eigenvalue according to the feature vector, and generating the first feature matrix according to the eigenvalue.

[0070] Specifically, obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table, where the feature vector refers to the M dimensions corresponding to the event, that is, the header of the data table, obtaining the specific value corresponding to each dimension of each event as the eigenvalue, and generating the first feature matrix according to the eigenvalue.

[0071] In some of these embodiments, obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table includes: obtaining the feature vector corresponding to the event to be recognized from the adjusted data table according to the association relationship between the event to be recognized and other events.

[0072] In some of these embodiments, the fraud risk recognition method based on a graph neural network further includes: obtaining the second association relationship adjacency matrix of the training event according to the data table; obtaining the second feature matrix of the training event according to the data table and the second association relationship adjacency matrix; and training the graph neural network by using the second association relationship adjacency matrix and the second feature matrix.

[0073] In some of these embodiments, training the graph neural network by using the second association relationship adjacency matrix and the second feature matrix includes: inputting the second association relationship adjacency matrix of the training event and the second feature matrix of the training event into the graph neural network to obtain a predicted value; constructing a loss function of the graph neural network according to the predicted value and the label value of the training event; and adjusting the parameters of the graph neural network according to the loss function.

[0074] The embodiments of the present application will be described and illustrated below through specific examples.

[0075] Figure 4 is a flowchart of the fraud risk identification method based on the graph neural network in this specific embodiment. As Figure 4 shown, the fraud risk identification method based on the graph neural network includes the following steps:

[0076] Step S410: Extract the historical case association relationship adjacency matrix from the historical cases according to the association relationship graph; the association relationship graph is obtained from the data table structure in the structured database.

[0077] Specifically, the historical case structured database is used to store information related to historical case auto insurance claims. The data in the relational database is divided according to the actual meaning of the business scenario, that is, the association relationship between case numbers is extracted according to the custom auto insurance anti-fraud association relationship to obtain the case association relationship adjacency matrix, and the adjacency matrix is used to represent the connection relationship between nodes in the association relationship graph. Exemplarily, the data table in the structured database is shown in Table 1.

[0078] Table 1 Data table in the structured database

[0079]

[0080] Step S420: According to the sequence of the obtained historical case association relationship adjacency matrix, combined with the historical case database, obtain the historical case association relationship adjacency matrix and corresponding features.

[0081] Specifically, the historical case association relationship adjacency matrix is an n*n 0-1 matrix, where n is the number of cases. The value of the i-th row and j-th column being 0 indicates that there is no associated relationship edge between case i and case j. Similarly, the value of the i-th row and j-th column being 1 indicates that there is an associated relationship edge between case i and case j. As shown in Table 1, each case has information data in 5 dimensions: PlateNumber, ContactPersonTelephoneNumber, ContactPersonName, AccidentProvince, and AccidentCity. Each dimension corresponds to an n*n adjacency matrix, and the rows and columns of each adjacency matrix represent case numbers. If the order of cases in the obtained n*n adjacency matrix is 00, 02, 01, 04, 03*00, 02, 01, 04, 03, then the order of events in the data table needs to be adjusted to 00, 02, 01, 04, 03, and the corresponding features are extracted according to the adjusted data table. The corresponding features of the historical case association relationship adjacency matrix are the valid fields extracted from the table. For example, the features extracted from Table 1 are the values corresponding to PlateNumber, ContactPersonTelephoneNumber, ContactPersonName, AccidentProvince, and AccidentCity. The corresponding feature can be a feature matrix generated from the values corresponding to PlateNumber, ContactPersonTelephoneNumber, ContactPersonName, AccidentProvince, and AccidentCity extracted from Table 1. This feature matrix is an n*m matrix, where m is the number of dimensions of the data in the data table. As there are 5 dimensions of information data in Table 1, m is 5 at this time. The relevant fields are sorted in sequence according to the case numbers in the case association relationship adjacency matrix. If there are multiple fields representing the same attribute value, then the field with the largest amount of information needs to be selected as the attribute field. After final sorting, a database table is formed into the case association relationship adjacency matrix and the corresponding features.

[0082] Step S430, use the historical case association relationship adjacency matrix and the corresponding features to train the graph neural network model.

[0083] A graph neural network refers to a topological graph that uses vertices and edges to establish corresponding relationships in graph theory in mathematics, and at the same time uses graph adjacency matrix information and corresponding feature information.

[0084] The message passing definition of a basic graph neural network is shown in formula (1):

[0085]

[0086] Among them, A is the adjacency matrix, Denote the node representation matrix of the k-th layer in the graph neural network (each node corresponds to a row in the matrix), and is a trainable parameter matrix, and σ represents an element-wise non-linear function (such as tanh or ReLU).

[0087] The loss function of the graph neural network model is calculated using the cross-entropy formula, as shown in Equation (2):

[0088] H(p,q) = -∑ x (p(x)logq(x) + (1 - p(x))log(1 - q(x))) (2)

[0089] where x is the training case, p(x) is the true label, the value of p(x) is in [0,1], representing the fraud risk of the training case x. The larger the probability value, the greater the fraud risk. q(x) is the label predicted according to the above neural network.

[0090] The model updates the gradient using the stochastic gradient descent method, as shown in Equation (3):

[0091]

[0092] where x is the training case input to the model, y is the true label of case x, δ t is the set of all samples, the function f(x;θ) is the result calculated by the current model when the input data is x and the parameter is θ, the function L(y,f(x;θ)) is the loss function calculated by substituting the result of the current model data calculation and the true label of the case. This loss function can be calculated using Equation (2), f(x;θ t ) is the result calculated by the current model when the input data is x and the parameter is θ t in the case, the function L(y,f(x;θ t )) is the loss function calculated by substituting the result of the current model data calculation and the true label of the case. The θ in the next state t+1 is the θ in the previous state t , the average value of K sample sampling points is obtained by taking the partial derivative of θ in the previous state with respect to the prediction result y of the current time, and then multiplied by the gradient step size α.

[0093] Furthermore, update the weights of the neurons. Substitute the neuron weights W into θ in the gradient descent formula to obtain Equation (4). Here, the neuron weights W are or

[0094]

[0095] After passing the case association relationship adjacency matrix, corresponding features, and the label of whether the case is fraudulent into the training code of the graph neural network model, the graph neural network model can be trained. The training code of the graph neural network model is rewritten according to the specific data situation of the auto insurance claim settlement business for the graph neural network model.

[0096] Step S440: Extract the historical case association relationship adjacency matrix for the case to be predicted according to the association relationship graph; the association relationship graph is obtained from the data table structure in the structured database.

[0097] Step S450: According to the sequence of the obtained association relationship adjacency matrix of the case to be predicted, combine it with the dataset of the case to be predicted to obtain the association relationship adjacency matrix and corresponding features of the case to be predicted.

[0098] Specifically, the dataset of the case to be predicted here can be a data table in the structured database.

[0099] Step S460: Use the trained graph neural network model in Step S430 to predict the fraud risk probability of the case, and obtain the risk value of the case to be predicted for auto insurance.

[0100] Specifically, according to the predicted risk value of the case, retrieve the case from the relational database, and return the case number and risk value of the case to the user.

[0101] In this embodiment, by obtaining the association relationship adjacency matrix according to the historical data in the data table, and obtaining the corresponding features of the event to be identified according to the association relationship adjacency matrix, according to the association relationship adjacency matrix and corresponding features, obtaining the fraud probability of the event to be identified through the graph neural network, and performing risk identification according to the fraud probability, thus solving the problem that the existing fraud risk identification method in the field of auto insurance claim settlement does not associate the historical data in the database, resulting in inaccurate identification results.

[0102] A specific embodiment is given below. In this embodiment, there are 9,175 cases in the auto insurance dataset, of which 2,129 cases are marked as fraudulent cases, and the case fraud rate is 23.20%.

[0103] When processing the historical case dataset, the reported call feature in the relational database is used, and the cases with the same reported call are connected by edges. In the adjacency matrix, the edge relationship corresponding to the case is set to 1.

[0104] In the embodiment, a 2-layer graph convolutional neural network model is adopted. Each hidden layer uses 100 nodes, the dropout rate is set to 0.5, the learning rate is 0.001, and it is trained 100,000 rounds at a time. Use a part of the historical case association relationship adjacency matrix and corresponding features obtained by splitting as the training set for model training, and the remaining cases are used as the test set for model testing.

[0105] When comparing with a general neural network, a two-layer neural network is used, with 100 hidden layer nodes in each layer, a dropout rate, and a learning rate of 0.001. The number of training epochs is the same as that of the graph convolutional neural network model. Considering that the general neural network cannot use graph association relationship information, only the graph nodes of the general neural network are set as a self-correlated diagonal matrix, and the same training set is used for model training, and the same test set is used to test the model metrics.

[0106] In actual business scenarios, insurance companies will use models to evaluate the risk value of individual cases, and for cases with higher risks, manual investigations are adopted to determine whether the cases are fraud cases. Considering that there is a certain cost for each manual investigation of a case, insurance companies hope that the fraud rate of the cases investigated manually is as high as possible. Generally, the case sampling rate of insurance companies is between 1% and 5%.

[0107] In order to be able to evaluate the advantages and disadvantages of different algorithms, the concept of F1 value is proposed based on precision and recall to conduct an overall evaluation of precision and recall. The definition of F1 is as follows, as shown in formula (5):

[0108] F1 value = precision * recall * 2 / (precision + recall)

[0109]

[0110] As shown in Table 2, the AUC value, accuracy, precision, recall, and F1 value of the graph neural network. It can be seen from Table 2 that when only using the corresponding case features and not using the graph adjacency matrix, the F1 value obtained in the neural network model is only 0.085. After using the graph adjacency matrix, the F1 value is improved to 0.243, and the improvement is relatively obvious.

[0111] Table 2 Evaluation index table of graph neural network model and neural network

[0112] Model Name AUC Value Accuracy Precision Recall F1 Score Graph Neural Network 0.555 0.694 0.646 0.150 0.243 Neural Network 0.515 0.778 0.442 0.047 0.085

[0113] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0114] In this embodiment, a fraud risk identification device based on a graph neural network is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0115] Figure 5 is a structural block diagram of a fraud risk identification device based on a graph neural network according to an embodiment of the present application. As Figure 5 shown, the device includes:

[0116] A first acquisition module 510, configured to obtain a first association relationship adjacency matrix of an event to be identified according to a data table; the data table includes data of the event to be identified, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be identified and other events in the data table;

[0117] A second acquisition module 520, configured to obtain a first feature matrix of the event to be identified according to the data table and the first association relationship adjacency matrix; the first feature matrix is used to identify the data of the events in the data table;

[0118] A third acquisition module 530, configured to obtain a fraud probability of the event to be identified according to the first association relationship adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of an event;

[0119] A determination module 540, configured to determine the fraud risk of the event to be identified according to the fraud probability.

[0120] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.

[0121] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0122] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.

[0123] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:

[0124] S1. Obtain the first associated relationship adjacency matrix of the event to be recognized according to the data table; the data table includes the data of the event to be recognized, and the first associated relationship adjacency matrix is used to identify the associated relationship between the event to be recognized and other events in the data table;

[0125] S2. Obtain the first feature matrix of the event to be recognized according to the data table and the first associated relationship adjacency matrix; the first feature matrix is used to identify the data of the events in the data table;

[0126] S3. Obtain the fraud probability of the event to be recognized according to the first associated relationship adjacency matrix, the first feature matrix, and the pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of the event;

[0127] S4. Determine the fraud risk of the event to be recognized according to the fraud probability.

[0128] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated in this embodiment.

[0129] In addition, in combination with the fraud risk recognition method based on the graph neural network provided in the above embodiments, a storage medium can also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by a processor, the steps of any one of the fraud risk recognition methods based on the graph neural network in the above embodiments are implemented.

[0130] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present application.

[0131] Obviously, the drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations according to these drawings without creative work. In addition, it can be understood that although the work done during the development process here may be complex and time-consuming, for those of ordinary skill in the art, some design, manufacturing, or production changes made according to the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.

[0132] The term "embodiment" in this application means that the specific features, structures or characteristics described in connection with an embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.

[0133] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A fraud risk identification method based on graph neural network, characterized in that, The method includes: Obtaining an association relationship graph of the event to be recognized according to a data table; the association relationship graph is used to identify the association relationship between the event to be recognized and other events in the data table; generating a first association relationship adjacency matrix of the event to be recognized according to the association relationship graph; the data table includes data of the event to be recognized, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be recognized and other events in the data table; Adjusting the record data of the data table according to the event order in the first association relationship adjacency matrix; obtaining a first feature matrix of the event to be recognized according to the adjusted data table; the first feature matrix is used to identify the data of the events in the data table; Obtaining the fraud probability of the event to be recognized according to the first association relationship adjacency matrix, the first feature matrix and a pre-trained graph neural network; the graph neural network is used to obtain the fraud probability of an event; Determining the fraud risk of the event to be recognized according to the fraud probability.

2. The fraud risk identification method based on a graph neural network according to claim 1, characterized in that The obtaining the first feature matrix of the event to be recognized according to the adjusted data table includes: Obtaining a feature vector corresponding to the event to be recognized according to the adjusted data table, obtaining a corresponding eigenvalue according to the feature vector, and generating the first feature matrix according to the eigenvalue.

3. The fraud risk identification method based on a graph neural network according to claim 2, wherein The obtaining the feature vector corresponding to the event to be recognized according to the adjusted data table includes: Obtaining the feature vector corresponding to the event to be recognized from the adjusted data table according to the association relationship between the event to be recognized and other events.

4. The fraud risk identification method based on graph neural network according to claim 1, characterized in that The method further includes: Obtaining a second association relationship adjacency matrix of a training event according to the data table; Obtaining a second feature matrix of the training event according to the data table and the second association relationship adjacency matrix; Training the graph neural network by using the second association relationship adjacency matrix and the second feature matrix.

5. The fraud risk identification method based on a graph neural network according to claim 4, wherein The training the graph neural network by using the second association relationship adjacency matrix and the second feature matrix includes: Inputting the second association relationship adjacency matrix of the training event and the second feature matrix of the training event into the graph neural network to obtain a predicted value; Constructing a loss function of the graph neural network according to the predicted value and the label value of the training event; Adjusting the parameters of the graph neural network according to the loss function.

6. A fraud risk identification device based on a graph neural network, characterized in that, The device includes: A first obtaining module, configured to obtain an association relationship graph of the event to be recognized according to a data table; the association relationship graph is used to identify the association relationship between the event to be recognized and other events in the data table; generating a first association relationship adjacency matrix of the event to be recognized according to the association relationship graph; the data table includes data of the event to be recognized, and the first association relationship adjacency matrix is used to identify the association relationship between the event to be recognized and other events in the data table; A second acquisition module, configured to adjust the record data in the data table according to the event order in the first correlation relationship adjacency matrix; and acquire a first feature matrix of the event to be recognized according to the adjusted data table; the first feature matrix is used to identify the data of the events in the data table. A third acquisition module, configured to acquire the fraud probability of the event to be recognized according to the first correlation relationship adjacency matrix, the first feature matrix, and a pre-trained graph neural network; the graph neural network is used to acquire the fraud probability of an event. A determination module, configured to determine the fraud risk of the event to be recognized according to the fraud probability.

7. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the graph neural network-based fraud risk identification method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the graph neural network-based fraud risk identification method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Risk identification method and device for vehicle insurance claim settlement case, equipment and storage medium

    CN109919783A

  • Anti-fraud method and system based on heterogeneous graph neural network, and computer equipment

    CN114579991A