Transaction risk prediction method, related device and medium

Through training and feature reconstruction based on the logistic regression model, the problems of inaccurate risk prediction and insufficient real-time performance in existing technologies are solved, and more efficient risk prediction accuracy and real-time performance are achieved.

CN120706864APending Publication Date: 2025-09-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410357777.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The risk prediction methods in the existing technology have problems with inaccurate predictions and lack of real-time performance in image, natural language processing and risk warning scenarios. The established rule combinations are single, the neural network model is difficult to train, and the computational complexity is high.

Method used

The logistic regression model is trained using sample transaction risk scores and sample transaction labels based on multiple sample transactions. The sample reconstruction coefficient vector is predicted using reference risk and security transaction features. The target transaction risk score is calculated through the reconstruction function and prediction network to achieve risk prediction model training and transaction processing.

Benefits of technology

The accuracy and real-time performance of risk prediction are improved. The training process of the logistic regression model is simple and the computational complexity is low. It can effectively eliminate invalid feature interference and improve fitting and prediction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706864A_ABST
    Figure CN120706864A_ABST
Patent Text Reader

Abstract

The invention provides a transaction risk prediction method, a related device and a medium. The method comprises the steps of obtaining target transaction characteristics of a target transaction; inputting the target transaction characteristics of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction, and performing transaction processing on the target transaction based on the target transaction risk score; the risk prediction model is obtained by training a logistic regression model based on comparison of a sample transaction risk score and a sample transaction label of the sample transaction; the sample transaction risk score is obtained by inputting a sample reconstruction coefficient vector of the sample transaction into a logistic regression model; the sample reconstruction coefficient vector is predicted according to the sample transaction associated data of the sample transaction, preset reference risk transaction characteristics and reference security transaction characteristics. The risk prediction accuracy of the transaction can be improved, and the real-time performance of risk prediction can be improved at the same time. The method can be applied to various scenes such as big data, artificial intelligence and cloud technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and in particular to a transaction risk prediction method, related devices, and media. Background Art

[0002] Currently, in application scenarios such as image processing, natural language processing, and risk warning, it is often necessary to control the risks of transactions generated by multiple business processes, such as resource interaction and information verification, to improve transaction security. Risk prediction methods in related technologies often rely on established rule combinations or neural network models. However, risk prediction based on established rule combinations often suffers from inaccurate predictions due to the relatively simple set of rules. Risk prediction based on neural network models often fails to meet real-time requirements due to training difficulties and high computational complexity. Summary of the Invention

[0003] The embodiments of the present disclosure provide a transaction risk prediction method, related devices, and media, which can improve the accuracy of transaction risk prediction while improving the real-time performance of risk prediction.

[0004] According to one aspect of the present disclosure, a transaction risk prediction method is provided, the method comprising:

[0005] Obtain target transaction characteristics of the target transaction;

[0006] Inputting the target transaction characteristics of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction, and performing transaction processing on the target transaction based on the target transaction risk score;

[0007] Among them, the risk prediction model is obtained by training a preset logistic regression model based on the comparison of sample transaction risk scores of multiple sample transactions and sample transaction labels of the sample transactions; the sample transaction risk score is obtained by inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model; the sample reconstruction coefficient vector is predicted based on the sample transaction association data of the sample transaction, and the pre-set reference risk transaction characteristics and reference security transaction characteristics.

[0008] According to one aspect of the present disclosure, a transaction risk prediction device is provided, the device comprising:

[0009] an acquisition unit, configured to acquire target transaction characteristics of a target transaction;

[0010] a prediction unit, configured to input a target transaction feature of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction, and perform transaction processing on the target transaction based on the target transaction risk score;

[0011] Among them, the risk prediction model is obtained by training a preset logistic regression model based on the comparison of sample transaction risk scores of multiple sample transactions and sample transaction labels of the sample transactions; the sample transaction risk score is obtained by inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model; the sample reconstruction coefficient vector is predicted based on the sample transaction association data of the sample transaction, and the pre-set reference risk transaction characteristics and reference security transaction characteristics.

[0012] Optionally, the risk prediction model includes a reconstruction function and a prediction network;

[0013] The prediction unit is specifically configured to:

[0014] Calculating a target reconstruction coefficient vector of the target transaction based on the reference risk transaction feature, the reference safety transaction feature, and the target transaction feature by the reconstruction function;

[0015] The target reconstruction coefficient vector is input into the prediction network for probability calculation to obtain the target transaction risk score.

[0016] Optionally, the performing transaction processing on the target transaction based on the target transaction risk score includes:

[0017] Invoking a plurality of predetermined candidate strategies and triggering conditions corresponding to each of the predetermined candidate strategies;

[0018] If it is determined that the target transaction risk score and the target transaction characteristics meet the triggering condition of a target policy among the plurality of predetermined candidate policies, transaction processing is performed on the target transaction based on the target policy.

[0019] Optionally, the transaction risk prediction device further includes a training unit, which specifically includes:

[0020] a sample acquisition unit, configured to acquire sample transaction association data and sample transaction tags of a plurality of the sample transactions;

[0021] a mapping unit, configured to obtain, for each sample transaction, a sample transaction feature of the sample transaction based on a mapping result of the sample transaction-related data;

[0022] a determining unit, configured to determine, for each of the sample transactions, a sample reconstruction coefficient vector of the sample transaction based on the sample transaction feature, the reference risk transaction feature, and the reference safety transaction feature;

[0023] an input unit, configured to input the sample reconstruction coefficient vector into the logistic regression model for each sample transaction, and output a sample transaction risk score of the sample transaction;

[0024] An adjustment unit is used to adjust the model parameters of the logistic regression model based on a comparison of the sample transaction risk scores of the plurality of sample transactions and the sample transaction labels to obtain the risk prediction model.

[0025] Optionally, the sample transaction associated data includes sample transaction content of the sample transaction and a sample transaction associated object;

[0026] The mapping unit is used for:

[0027] Obtaining a sample transaction content feature of the sample transaction based on a mapping result of the sample transaction content;

[0028] Obtaining a sample transaction-associated object feature of the sample transaction based on a mapping result of the sample transaction-associated object;

[0029] The sample transaction content feature and the sample transaction associated object feature are integrated into the sample transaction feature.

[0030] Optionally, the determining unit is configured to:

[0031] generating a predetermined transaction feature matrix based on the reference risk transaction feature and the reference safety transaction feature;

[0032] Determining an objective function based on the predetermined transaction feature matrix and a predetermined regularization coefficient;

[0033] For each of the sample transactions, the sample transaction characteristics are substituted into the objective function, and the regularization term parameter that minimizes the output result of the objective function is determined as the sample reconstruction coefficient vector of the sample transaction.

[0034] Optionally, the logistic regression model includes an activation function, a first parameter, and a second parameter;

[0035] The input unit is used to:

[0036] Inputting the sample reconstruction coefficient vector into the logistic regression model, and obtaining a first calculation result based on a transposed result of the first parameter and a product of the sample reconstruction coefficient vector;

[0037] Probability scoring is performed on the sum of the first calculation result and the second parameter based on the activation function to obtain a sample transaction risk score of the sample transaction.

[0038] Optionally, the adjustment unit is configured to:

[0039] For each of the sample transactions, calculating a sub-loss function based on the sample transaction risk score and the sample transaction label;

[0040] Obtaining a total loss function based on the sum of the multiple sub-loss functions;

[0041] The model parameters of the logistic regression model are adjusted based on the total loss function to obtain the risk prediction model.

[0042] Optionally, the sample transactions include positive sample transactions and negative sample transactions;

[0043] The calculating of a sub-loss function for each of the sample transactions based on the sample transaction risk score and the sample transaction label includes:

[0044] If it is determined that the sample transaction is a positive sample transaction, determining the sub-loss function based on the negative logarithm of the risk score of the sample transaction;

[0045] If it is determined that the sample transaction is a negative sample transaction, the sub-loss function is determined based on the negative logarithm of the difference between 1 and the risk score of the sample transaction.

[0046] Optionally, the transaction risk prediction device further includes a first testing unit, wherein the first testing unit is specifically configured to:

[0047] Obtaining first transaction features and first transaction tags of a plurality of first test transactions;

[0048] For each of the first test transactions, inputting the first transaction feature of the first test transaction into the risk prediction model to obtain a first transaction risk score of the first test transaction;

[0049] Based on the first transaction risk scores of the plurality of first test transactions and the first transaction labels, a first model verification is performed on the risk prediction model, so as to update the model parameters of the risk prediction model when it is determined that the first model verification fails.

[0050] Optionally, the transaction risk prediction device further includes a second testing unit, and the second testing unit is specifically configured to:

[0051] If it is determined that the first model verification passes, obtaining second transaction features and second transaction tags of a plurality of second test transactions, wherein a transaction time difference between the second test transaction and the first test transaction satisfies a predetermined condition;

[0052] For each second test transaction, inputting the second transaction feature of the second test transaction into the risk prediction model to obtain a second transaction risk score of the second test transaction;

[0053] Based on the second transaction risk scores of multiple second test transactions and the second transaction labels, the risk prediction model is subjected to a second model verification, so as to update the model parameters of the risk prediction model when it is determined that the second model verification fails, or to deploy the risk prediction model online when the second model verification passes.

[0054] Optionally, performing a second model validation on the risk prediction model based on the second transaction risk scores of the plurality of second test transactions and the second transaction labels includes:

[0055] determining a confusion matrix of the risk prediction model based on the second transaction risk scores of the plurality of second test transactions, the second transaction labels, and a predetermined threshold;

[0056] According to a first validation rule, based on the confusion matrix, performing a first sub-validation on the risk prediction model;

[0057] According to the second verification rule, based on the confusion matrix, a second sub-verification is performed on the risk prediction model.

[0058] Optionally, performing a first sub-validation on the risk prediction model according to the first validation rule and based on the confusion matrix includes:

[0059] Determining, based on the confusion matrix, an information ratio, a maximum cumulative transaction risk score difference, and a feature discrimination accuracy of the risk prediction model;

[0060] Based on the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy, a first sub-verification is performed on the risk prediction model.

[0061] Optionally, performing a second sub-validation on the risk prediction model according to a second validation rule and based on the confusion matrix includes:

[0062] Determining, based on the confusion matrix, a first number of second test transactions whose second transaction risk scores are greater than or equal to a predetermined threshold, a second number of second test transactions whose second transaction label is 1, and a third number of second test transactions whose second transaction risk scores are greater than or equal to the predetermined threshold and whose second transaction label is 1;

[0063] Determining, according to the second verification rule, a transaction coverage of the risk prediction model based on a quotient of the third number and the second number, and determining a strategy cost-effectiveness of the risk prediction model based on a quotient of the first number and the third number;

[0064] A second sub-verification is performed on the risk prediction model based on the transaction coverage and the strategy cost-effectiveness.

[0065] Optionally, the predetermined threshold is determined by:

[0066] Determining the current application scenario type of the risk prediction model;

[0067] The predetermined threshold is determined from a plurality of candidate thresholds based on the current application scenario type.

[0068] Optionally, the predetermined regularization coefficient is determined by:

[0069] Determining the sparsity of the sample reconstruction coefficient vector;

[0070] The predetermined regularization coefficient is determined from a plurality of candidate regularization coefficients based on the sparsity.

[0071] Optionally, there are multiple risk prediction models, and each candidate transaction type corresponds to one risk prediction model;

[0072] The prediction unit is used to:

[0073] Determining a target transaction type of the target transaction from a plurality of candidate transaction types based on the target transaction characteristics;

[0074] The target transaction features are input into the risk prediction model corresponding to the target transaction type to obtain a target transaction risk score of the target transaction.

[0075] Optionally, the predetermined condition is determined by:

[0076] Determining a training accuracy of the risk prediction model and a total number of predetermined transactions of the second test transaction;

[0077] Based on the training accuracy and the total number of predetermined transactions, an extreme value of a transaction time difference between the second test transaction and the first test transaction is determined, and the predetermined condition is determined based on the extreme value of the time difference.

[0078] Optionally, the transaction risk prediction device further includes a collection unit, which is specifically configured to:

[0079] Determining data collection conditions for the risk prediction model;

[0080] If it is determined that the model training and model operation of the risk prediction model meet the data collection conditions, the log data of the risk prediction model is obtained to update the model based on the log data.

[0081] According to one aspect of the present disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the transaction risk prediction method as described above when executing the computer program.

[0082] According to one aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the transaction risk prediction method as described above is implemented.

[0083] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the transaction risk prediction method as described above.

[0084] In the disclosed embodiments, the risk prediction model is obtained by training a logistic regression model based on the sample transaction risk scores and sample transaction labels of the sample transactions. Compared to neural network models in related art, the training process of the logistic regression model is relatively simple, and the computational complexity of the logistic regression model is relatively low, thus meeting the real-time requirements of risk prediction. Furthermore, when training the logistic regression model, a sample reconstruction coefficient vector for the sample transaction is predicted based on reference risk transaction features, reference security transaction features, and sample transaction association data of the sample transaction. This extracts effective features of the sample transaction and eliminates interference from invalid features. Furthermore, this approach considers inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model to obtain the sample transaction risk score. The use of the sample reconstruction coefficient vector imbues the logistic regression model with a certain degree of nonlinearity, which helps improve the fitting and predictive capabilities of the risk prediction model. Finally, the acquired target transaction features are input into the risk prediction model for risk prediction, thereby outputting a target transaction risk score for the target transaction based on the risk prediction model. Transaction processing is then performed based on the target transaction risk score. Compared to risk prediction methods based on rule combinations in related art, this approach can effectively improve the accuracy of risk prediction.

[0085] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present disclosure. The purposes and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The accompanying drawings are used to provide a further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0087] Figure 1 is a system architecture diagram of a system to which the transaction risk prediction method according to an embodiment of the present disclosure is applied;

[0088] Figure 2A-2C A schematic diagram showing the application of a transaction risk prediction method in a resource interaction scenario according to an embodiment of the present disclosure is shown;

[0089] Figure 3 is a flowchart of a transaction risk prediction method according to an embodiment of the present disclosure;

[0090] Figure 4 This is a flow chart of an embodiment of the present disclosure for inputting target transaction features into a risk prediction model to obtain a target transaction risk score;

[0091] Figure 5This is a schematic diagram of an implementation process of inputting target transaction features into a risk prediction model to obtain a target transaction risk score according to an embodiment of the present disclosure;

[0092] Figure 6 is a flowchart of another embodiment of the present disclosure for inputting target transaction features into a risk prediction model to obtain a target transaction risk score;

[0093] Figure 7 is a flowchart of processing a target transaction based on a target transaction risk score according to an embodiment of the present disclosure;

[0094] Figure 8 is a flowchart of training a risk prediction model according to an embodiment of the present disclosure;

[0095] Figure 9 is a flow chart of obtaining sample transaction features of a sample transaction according to an embodiment of the present disclosure;

[0096] Figure 10 is a flowchart of determining a sample reconstruction coefficient vector of a sample transaction according to an embodiment of the present disclosure;

[0097] Figure 11 is a flow chart of outputting a sample transaction risk score of a sample transaction according to one embodiment of the present disclosure;

[0098] Figure 12 is a flow chart of adjusting model parameters of a logistic regression model according to one embodiment of the present disclosure;

[0099] Figure 13 Schematic diagram of the overall implementation process of model training and model application according to one embodiment of the present disclosure;

[0100] Figure 14 is a flow chart of performing model evaluation on a risk prediction model according to one embodiment of the present disclosure;

[0101] Figure 15 is a flow chart of performing model evaluation on a risk prediction model according to another embodiment of the present disclosure;

[0102] Figure 16 is a flowchart of determining predetermined conditions to be satisfied by a first test transaction and a second test transaction according to an embodiment of the present disclosure;

[0103] Figure 17 is a flow chart of performing a second model validation on a risk prediction model according to one embodiment of the present disclosure;

[0104] Figure 18 is a flowchart of performing a first sub-validation of a risk prediction model according to one embodiment of the present disclosure;

[0105] Figure 19 is a flowchart of performing a second sub-validation of a risk prediction model according to one embodiment of the present disclosure;

[0106] Figure 20 is a flow chart of data collection for model training and operation according to one embodiment of the present disclosure;

[0107] Figure 21 1 is a schematic diagram of an implementation process of data collection for model training and operation according to an embodiment of the present disclosure;

[0108] Figures 22A-22B is a schematic diagram of implementation details of a transaction risk prediction method according to an embodiment of the present disclosure;

[0109] Figure 23 is a module diagram of a transaction risk prediction device according to an embodiment of the present disclosure;

[0110] Figure 24 is a terminal structure diagram of a transaction risk prediction method according to an embodiment of the present disclosure;

[0111] Figure 25 It is a server structure diagram of a transaction risk prediction method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0112] In order to make the purpose, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not intended to limit the present disclosure.

[0113] Before further explaining the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:

[0114] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning. With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0115] Currently, in application scenarios such as image processing, natural language processing, and risk warning, it is often necessary to control the risks of transactions generated by multiple business processes, such as resource interaction and information verification, to improve transaction security. Risk prediction methods in related technologies often rely on established rule combinations or neural network models. However, risk prediction based on established rule combinations often suffers from inaccurate predictions due to the relatively simple set of rules. Risk prediction based on neural network models often fails to meet real-time requirements due to training difficulties and high computational complexity.

[0116] System architecture and scenario description of the application of the embodiments of the present disclosure

[0117] Figure 1 This is a system architecture diagram for the transaction risk prediction method according to an embodiment of the present disclosure, which includes an object terminal 140, the Internet 130, a gateway 120, a transaction processing platform server 110, and a risk control policy library 150.

[0118] The target terminal 140 includes various forms, such as desktop computers, laptops, PDAs (personal digital assistants), mobile phones, in-vehicle terminals, home theater terminals, and dedicated terminals. Furthermore, it can be a single device or a collection of multiple devices. The target terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data. The target terminal 140 includes a transaction processing system, which allows targets to submit various types of transactions and request the execution of transaction processing for submitted transactions.

[0119] The transaction processing platform server 110 refers to a computer system that can provide certain services to the object terminal 140. Compared with the ordinary object terminal 140, the transaction processing platform server 110 has higher requirements in terms of stability, security, performance, etc. The transaction processing platform server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), etc. The server 110 includes various types of services, wherein the implementation of each service of the transaction processing platform server 110 is often associated with some intermediate databases or storage media. The transaction processing platform server 110 is used to identify the risks of transactions submitted by the object using the trained risk prediction model, so as to process the submitted transactions according to the transaction risk size of each transaction and the various risk control policies stored in the risk control policy library 150, so as to reduce the risks caused by executing unsafe transactions.

[0120] The gateway 120 is also known as a gateway or protocol converter. It implements network interconnection at the transport layer and is a computer system or device that performs a conversion function. It acts as a translator between two systems using different communication protocols, data formats, languages, or even completely different architectures. It can also provide filtering and security functions. Messages sent by the client terminal 140 to the transaction processing platform server 110 are sent to the corresponding server via the gateway 120. Messages sent by the transaction processing platform server 110 to the client terminal 140 are also sent to the corresponding client terminal 140 via the gateway 120.

[0121] The embodiments of the present disclosure can be applied in various scenarios, such as Figure 2A-2C The resource interaction scenarios shown, etc.

[0122] like Figure 2AAs shown, object A triggers a resource interaction process on the transaction processing platform of object terminal 140. The content page corresponding to the resource interaction process displays a prompt field "Please enter the following information" and provides input fields for the resource recipient address and interaction resource information. Based on this, object A enters "xqrrwr8978" in the input field for the resource recipient address and "100 virtual resources K" in the input field for interaction resource information. Then, object A clicks the "OK" button to submit a resource interaction transaction, exchanging its own 100 virtual resources K with the object whose terminal address is "xqrrwr8978."

[0123] like Figure 2B As shown, after object A clicks the "OK" button, the transaction processing platform will use the trained risk prediction model to perform risk prediction on the resource interaction transaction submitted by object A. At this time, a prompt field "Detecting the risk of this resource interaction transaction, please wait patiently..." is displayed on the content page of the transaction processing platform.

[0124] like Figure 2C As shown, the transaction processing platform uses the trained risk prediction model to perform risk detection on the resource interaction transaction submitted by object A. The content page of the transaction processing platform displays a prompt field stating, "The risk score of this resource interaction transaction is 90. The resource recipient presents a high risk. This resource interaction transaction has been automatically intercepted and the resource interaction has failed." This indicates that the resource interaction transaction submitted by the object is risky. Interacting virtual resource K with the object with the terminal address "xqrrwr8978" will make the resource unsafe. Based on this, object A clicks the "OK" button to confirm that the resource interaction transaction will not be executed.

[0125] General description of the embodiments of the present disclosure

[0126] According to one embodiment of the present disclosure, a transaction risk prediction method is provided.

[0127] This transaction risk prediction method is generally used in business scenarios with high security and real-time requirements, such as Figure 2A-2C The embodiment of the present disclosure provides a solution for training a logistic regression model into a risk prediction model based on the LASSO algorithm and performing transaction risk prediction based on the risk prediction model, which can improve the accuracy of transaction risk prediction while improving the real-time performance of risk prediction.

[0128] like Figure 3 As shown, a transaction risk prediction method according to an embodiment of the present disclosure may include:

[0129] Step 310: Obtain target transaction characteristics of the target transaction;

[0130] Step 320: Input the target transaction characteristics of the target transaction into the risk prediction model to obtain a target transaction risk score of the target transaction, so as to perform transaction processing on the target transaction based on the target transaction risk score.

[0131] Steps 310 - 320 are described in detail below.

[0132] In step 310 , target transaction features of the target transaction are obtained.

[0133] Target transactions refer to some online information interaction transactions, such as information exchange, resource interaction, identity authentication, and so on.

[0134] The target transaction characteristics refer to characteristic information such as the specific transaction content of the target transaction and the object characteristics of the transaction-related objects.

[0135] For example, for a resource interaction transaction occurring between two objects, the target transaction characteristics of the resource interaction transaction include the object characteristics of the two objects involved in the resource interaction transaction (for example, the regions where the two objects are located, the industries they are engaged in, etc.), the associated feature information between the two objects (for example, whether the two objects have any intersection in the past time, etc.), and the transaction content of the resource interaction transaction (for example, the type of virtual resources to be interacted, the number of resources, etc.).

[0136] In a specific implementation of this embodiment, with authorization, first, information such as the target transaction content and object characteristics of the transaction-related objects is obtained to form target transaction information of the target transaction. Next, the target transaction information of the target transaction is mapped from the data space to a preset vector space, and the representation of the target transaction information of the target transaction in the vector space is used as the target transaction characteristics of the target transaction, so that the target transaction information of the target transaction is converted into a form that can be received by the risk prediction model.

[0137] In step 320 , the target transaction characteristics of the target transaction are input into the risk prediction model to obtain a target transaction risk score of the target transaction, so as to perform transaction processing on the target transaction based on the target transaction risk score.

[0138] The target transaction risk score indicates the likelihood that the target transaction is risky. A higher target transaction risk score indicates a higher likelihood that the target transaction is risky. A lower target transaction risk score indicates a lower likelihood that the target transaction is risky.

[0139] Transaction processing refers to various forms of processing, such as ensuring the normal execution of the target transaction or intercepting the target transaction.

[0140] A risk prediction model refers to a model constructed using artificial intelligence technology that can be used to predict transaction risk. In the disclosed embodiments, the input to the risk prediction model is often the transaction characteristics of various transactions, and the output of the risk prediction model is the risk score of each transaction. The role of the risk prediction model is to autonomously determine whether a target transaction is a risky transaction based on the target transaction characteristics and output a target transaction risk score for the target transaction.

[0141] To save space, the specific implementation process of inputting the target transaction characteristics into the risk prediction model to obtain the target transaction risk score of the target transaction, and the specific implementation process of transaction processing the target transaction based on the target transaction risk score in the embodiment of the present disclosure will be described in detail below. No further details will be given here.

[0142] It should be noted that in the disclosed embodiments, the risk prediction model is obtained by training a preset logistic regression model based on a comparison of the sample transaction risk scores of multiple sample transactions and their sample transaction labels. The sample transaction risk scores are obtained by inputting the sample reconstruction coefficient vectors of the sample transactions into the logistic regression model. The sample reconstruction coefficient vectors are predicted based on the sample transaction association data of the sample transactions and pre-set reference risk transaction features and reference security transaction features.

[0143] The logistic regression model is a generalized linear regression analysis model that belongs to supervised learning in machine learning. It is mainly used to solve binary classification problems and multi-classification problems.

[0144] Sample transactions refer to transactions used to train the logistic regression model.

[0145] The sample transaction risk score is used to indicate the possibility that the sample transaction is risky.

[0146] The sample transaction tag is used to identify whether the sample transaction is truly a risky transaction. When the sample transaction is a risky transaction, the sample transaction tag is 1; when the sample transaction is a safe transaction (non-risky transaction), the sample transaction tag is 0.

[0147] The sample transaction association data is used to indicate the specific transaction content associated with the sample transaction, as well as object information of each object involved, and the like.

[0148] The sample reconstruction coefficient vector is used to indicate the result of feature selection on the transaction features of the sample transactions.

[0149] The preset reference risk transaction features indicate the transaction characteristics of reference risk transactions, while the reference safety transaction features indicate the transaction characteristics of reference safety transactions. To improve the model's ability to learn risk and safety transaction information, the ratio of the preset reference risk transactions to the preset reference safety transactions is fixed. For example, the ratio of reference risk transactions to reference safety transactions is set to 1:1.

[0150] To save space, the specific implementation process of training the preset logistic regression model and the specific implementation process of determining the sample reconstruction coefficient vector in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0151] Through the above steps 310-320, in the embodiment of the present disclosure, the risk prediction model is obtained by training the logistic regression model based on the sample transaction risk score of the sample transaction and the sample transaction label of the sample transaction. Compared with the neural network model of the related technology, the training process of the logistic regression model is relatively simple, and the computational complexity of the logistic regression model is relatively low, which can meet the real-time requirements of risk prediction. In addition, when training the logistic regression model, the sample reconstruction coefficient vector of the sample transaction is predicted based on the reference risk transaction characteristics, the reference security transaction characteristics, and the sample transaction association data of the sample transaction, thereby realizing the extraction of effective features of the sample transaction and eliminating the interference of invalid features. Furthermore, this method also considers inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model to obtain the sample transaction risk score. The sample reconstruction coefficient vector is used to give the logistic regression model a certain degree of nonlinearity, which is conducive to improving the fitting ability and prediction ability of the risk prediction model. Finally, the target transaction characteristics of the target transaction are input into the risk prediction model for risk prediction, so that the target transaction risk score of the target transaction is output based on the risk prediction model, and the transaction is processed according to the target transaction risk score. Compared with the risk prediction method based on rule combination in related technologies, this method can effectively improve the accuracy of risk prediction.

[0152] The above is a general description of steps 310 to 320. Since step 310 has been described in detail in the above general description, the following will describe step 320 and the specific implementation of training the risk prediction model in detail.

[0153] Detailed description of step 320

[0154] In step 320 , the target transaction characteristics of the target transaction are input into the risk prediction model to obtain a target transaction risk score of the target transaction, so as to perform transaction processing on the target transaction based on the target transaction risk score.

[0155] In a specific implementation of this embodiment, the risk prediction model of the disclosed embodiment includes a reconstruction function and a prediction network. The reconstruction function is used to perform feature screening on target transaction features to extract valid feature information from the target transaction features and eliminate interference from invalid feature information. The prediction network is used to score the target transaction based on the valid feature information extracted by the reconstruction function to determine the probability score of the target transaction being a risky transaction and the probability score of the target transaction being a safe transaction.

[0156] Please refer to Figure 4 In one embodiment, step 320 specifically includes but is not limited to the following steps 410-420:

[0157] Step 410: Calculate a target reconstruction coefficient vector of the target transaction based on the reference risk transaction characteristics, the reference safety transaction characteristics, and the target transaction characteristics through a reconstruction function;

[0158] Step 420: Input the target reconstruction coefficient vector into the prediction network for probability calculation to obtain the target transaction risk score.

[0159] Steps 410 - 420 are described in detail below.

[0160] In step 410 , a target reconstruction coefficient vector of the target transaction is calculated based on the reference risky transaction features, the reference safe transaction features, and the target transaction features through a reconstruction function.

[0161] In the specific implementation of this embodiment, first, a predetermined transaction feature matrix is ​​constructed based on the reference risk transaction features and the reference security transaction features, and the reference risk transaction features and the reference security transaction features are used as elements of the predetermined transaction feature matrix. Next, the predetermined transaction feature matrix and the predetermined regularization coefficient are substituted into the reconstruction function as function parameters, and the target transaction features are substituted into the reconstruction function as independent variables. Furthermore, with the goal of minimizing the output result of the reconstruction function, the reconstruction coefficient that minimizes the output result of the reconstruction function is used as the target reconstruction coefficient vector of the target transaction using the gradient descent method. The reconstruction function can be expressed as shown in formula (1):

[0162]

[0163] Wherein, M refers to the target transaction feature of the target transaction, A refers to the predetermined transaction feature matrix generated based on the reference risk transaction feature and the reference security transaction feature, m is the reconstruction coefficient, and λ is the predetermined regularization coefficient.

[0164] Once the predetermined transaction feature matrix and the predetermined regularization coefficient are determined, the functional form of the reconstruction function is fixed. When the target transaction feature M is input, m is solved using gradient descent. The reconstruction coefficient m that minimizes the output of the reconstruction function is used as the target reconstruction coefficient vector for the target transaction.

[0165] It should be noted that after the logistic regression model is trained into a risk prediction model, the predetermined transaction feature matrix and predetermined regularization coefficient of the reconstruction function in the risk prediction model are fixed.

[0166] In step 420, the target reconstruction coefficient vector is input into the prediction network for probability calculation to obtain the target transaction risk score.

[0167] In the specific implementation of this embodiment, the prediction network includes a nonlinear activation function sigmoid function for probability prediction. Specifically, first, the target reconstruction coefficient vector is input into the prediction network; then, the sigmoid function of the prediction network uses the target reconstruction coefficient vector as an independent variable and outputs a probability distribution. Finally, the target transaction risk score is determined based on the probability vector indicating that the target transaction is a risky transaction in the probability distribution. The process of performing probability calculation on the target reconstruction coefficient vector based on the sigmoid function can be expressed as shown in formula (2):

[0168] p=sigmoid(w T m+b) Formula (2)

[0169] Among them, p is the target transaction risk score of the target transaction, w and b are model parameters in the risk prediction model. After the logistic regression model is trained into a risk prediction model, w and b are fixed parameters. T Refers to the transposition operation of w. m is the target reconstruction coefficient vector of the target transaction.

[0170] like Figure 5 As shown, first, the target transaction features of the target transaction are input into the risk prediction model. Then, the reconstruction function in the risk prediction model reconstructs the target transaction features based on the predetermined transaction feature matrix and predetermined regularization coefficient generated based on the reference risk transaction features and the reference security transaction features to extract the effective feature information in the target transaction features and obtain the target reconstruction coefficient vector. Furthermore, the prediction network performs a probability calculation on the target reconstruction coefficient vector, and obtains that the probability that the target transaction is a risky transaction is 0.42, and the probability that the target transaction is a security transaction is 0.58. Based on this, the risk prediction model will output that the target transaction risk score of the target transaction is 42 points, and the target transaction is a security transaction. Further, based on the target transaction risk score of the target transaction and the predetermined candidate strategy in the risk control strategy library 150, the target transaction is processed.

[0171] The benefit of this embodiment is that it utilizes the reconstruction function of the risk prediction model to select target transaction features, effectively eliminating irrelevant feature information and noise, and improving the purity of the feature information used for risk prediction. Furthermore, by utilizing the activation function in the prediction network for probability calculation, the target transaction risk score can be determined with minimal computational effort, meeting both the accuracy and real-time requirements of transaction risk prediction.

[0172] Because transaction security requirements vary across different business scenarios, setting thresholds based on the same criteria to determine whether a target transaction is a risky transaction often results in inaccurate transaction predictions. Therefore, the disclosed embodiments provide a solution that sets a corresponding risk prediction model for each transaction type, meeting the transaction risk prediction requirements across different business scenarios and improving transaction prediction accuracy.

[0173] The embodiment of the present disclosure has multiple risk prediction models, each candidate transaction type corresponds to a risk prediction model, and different risk prediction models have different corresponding thresholds. For business scenarios with higher security requirements, the threshold corresponding to the risk prediction model will be smaller. As long as the target transaction risk score of the target transaction is greater than this threshold, the target transaction is considered to be a risky transaction. For business scenarios with lower security requirements, the threshold corresponding to the risk prediction model will be larger. Only when the target transaction risk score of the target transaction is greater than this threshold, the target transaction is considered to be a risky transaction.

[0174] Please refer to Figure 6 In one embodiment, step 320 specifically includes but is not limited to the following steps 610-620:

[0175] Step 610: Determine a target transaction type of the target transaction from multiple candidate transaction types based on the target transaction characteristics;

[0176] Step 620: Input the target transaction characteristics into the risk prediction model corresponding to the target transaction type to obtain a target transaction risk score for the target transaction.

[0177] Steps 610-620 are described in detail below.

[0178] In step 610 , based on the target transaction characteristics, a target transaction type of the target transaction is determined from a plurality of candidate transaction types.

[0179] In the specific implementation of this embodiment, first, with authorization, the preset feature rules corresponding to each candidate transaction type are invoked. The preset feature rules corresponding to each candidate transaction type indicate the conditions that the transaction features corresponding to the candidate transaction type must satisfy. Next, the target transaction features are compared with the preset feature rules corresponding to each candidate transaction type. Furthermore, the candidate transaction type indicated by the preset feature rules satisfied by the target transaction features is determined as the target transaction type of the target transaction.

[0180] In step 620 , the target transaction features are input into a risk prediction model corresponding to the target transaction type to obtain a target transaction risk score for the target transaction.

[0181] In the specific implementation of this embodiment, the specific implementation process of step 620 is similar to the specific implementation process of steps 410-420 described above, and will not be repeated for the sake of space.

[0182] The benefit of this embodiment is that, for each candidate transaction type, the conditions that the transaction characteristics corresponding to the candidate transaction type must meet are set, and the risk prediction model corresponding to the transaction type to which the target transaction characteristics belong is called to perform risk prediction. A corresponding risk prediction model is set for each transaction type, and different risk prediction models have different corresponding thresholds. For business scenarios with higher security requirements, the threshold is set to a lower value; for business scenarios with lower security requirements, the threshold corresponding to the risk prediction model is set to a higher value. This can meet the transaction risk prediction needs in different business scenarios, thereby improving the accuracy of transaction prediction.

[0183] Please refer to Figure 7 In one embodiment, the specific process of processing a target transaction based on the target transaction risk score includes but is not limited to the following steps 710-720:

[0184] Step 710: Call multiple predetermined candidate strategies and trigger conditions corresponding to each predetermined candidate strategy;

[0185] Step 720: If it is determined that the target transaction risk score and the target transaction characteristics meet the triggering conditions of a target policy among the multiple predetermined candidate policies, then perform transaction processing on the target transaction based on the target policy.

[0186] Steps 710 - 720 are described in detail below.

[0187] In step 710, a plurality of predetermined candidate policies and a trigger condition corresponding to each predetermined candidate policy are called.

[0188] The predetermined candidate strategy is used to indicate the specific processing method to be performed on each transaction. The predetermined candidate strategy includes, but is not limited to, executing the target transaction normally, intercepting the target transaction, sending a prompt message for the target transaction, and the like.

[0189] The triggering condition corresponding to the predetermined candidate policy is used to define under what circumstances the predetermined candidate policy can be triggered.

[0190] In the specific implementation of this embodiment, with authorization and permission, the server can call multiple predetermined candidate strategies for transaction risk prediction and the triggering conditions corresponding to each predetermined candidate strategy from the risk control strategy library 150.

[0191] In step 720 , if it is determined that the target transaction risk score and the target transaction characteristics meet the triggering conditions of a target policy among a plurality of predetermined candidate policies, transaction processing is performed on the target transaction based on the target policy.

[0192] The target policy refers to a predetermined candidate policy that matches the target transaction among multiple predetermined candidate policies.

[0193] In a specific implementation of this embodiment, the target transaction risk score and target transaction characteristics are first compared with the trigger conditions of each predetermined candidate policy to determine whether the target transaction risk score and target transaction characteristics meet the trigger conditions of a predetermined candidate policy. Next, if it is determined that the target transaction risk score and target transaction characteristics meet the trigger conditions of a predetermined candidate policy, the predetermined candidate policy corresponding to the trigger conditions met by the target transaction risk score and target transaction characteristics is determined as the target policy. Furthermore, transaction processing is performed on the target transaction based on the specific policy content of the target policy.

[0194] For example, the specific policy content of predetermined candidate policy A is to intercept transactions, and the triggering condition is that the risk score of the transaction is greater than 85 points, and the transaction characteristics indicate that there is no intersection between the objects associated with the transaction. The specific policy content of predetermined candidate policy B is to execute transactions normally, and the triggering condition is that the risk score of the transaction is greater than 60 points, and the transaction characteristics indicate that there are more than 10 interactions between the objects associated with the transaction. Based on this, when the target transaction risk score of target transaction K is 90 points, and there is no intersection between the multiple objects associated with target transaction K, the triggering condition of predetermined candidate policy A is met. Based on this, the target transaction K is intercepted according to the specific policy content of predetermined candidate policy A to prevent the execution of target transaction K.

[0195] The benefit of this embodiment is that after the risk prediction model outputs the target transaction risk score of the target transaction, it is combined with the target transaction risk score and the transaction characteristics of the target transaction to determine whether the target transaction meets the triggering conditions of the pre-deployed predetermined candidate policies, so that the target transaction can be processed according to the deployed multiple candidate policies, so that when a target transaction is at risk, the relevant objects can be reminded in a timely manner, or the transaction can be intercepted, etc., thereby improving the security of transaction execution and reducing the risk of resource loss caused by transaction execution.

[0196] Detailed description of the training risk prediction model of an embodiment of the present disclosure

[0197] Since risk prediction based on a set combination of rules often has poor stability and is prone to overfitting due to the relatively simple combination of rules, it will lead to inaccurate risk prediction of transactions; risk prediction based on neural network models often cannot meet real-time requirements due to training difficulties and high computational complexity, and is difficult to deploy online. Based on this, the embodiment of the present disclosure provides a solution for training a logistic regression model into a risk prediction model based on the LASSO algorithm, which can effectively improve the fitting ability and prediction ability of the risk prediction model, so that when the risk prediction model is used for actual transaction risk prediction, it can take into account the accuracy and real-time requirements of transaction risk prediction.

[0198] Please refer to Figure 8 In one embodiment, the process of training the risk prediction model includes but is not limited to the following steps 810-850:

[0199] Step 810: Obtain sample transaction association data and sample transaction tags of multiple sample transactions;

[0200] Step 820: For each sample transaction, based on the mapping result of the sample transaction associated data, obtain the sample transaction features of the sample transaction;

[0201] Step 830: Obtain sample transaction association data and sample transaction tags of multiple sample transactions;

[0202] Step 840: For each sample transaction, input the sample reconstruction coefficient vector into the logistic regression model, and output the sample transaction risk score of the sample transaction;

[0203] Step 850: Based on the comparison of the sample transaction risk scores and sample transaction labels of the plurality of sample transactions, the model parameters of the logistic regression model are adjusted to obtain a risk prediction model.

[0204] Steps 810-850 are described in detail below.

[0205] In step 810 , sample transaction association data and sample transaction tags of a plurality of sample transactions are obtained.

[0206] In the specific implementation of this embodiment, to improve data storage security, the server backend often maintains transaction data for each sample transaction through log files and other means, and stores different transaction contents in a transaction database. Based on this, upon obtaining authorization, a portion of transaction content is first extracted from the transaction database as a sample transaction. Next, the stored log file is retrieved from the server backend. Finally, the sample transaction association data for multiple sample transactions is read from the log file, and a sample transaction label is determined for each sample transaction based on whether the sample transaction is a risky transaction.

[0207] In step 820 , for each sample transaction, based on the mapping result of the sample transaction associated data, a sample transaction feature of the sample transaction is obtained.

[0208] The mapping result is used to indicate a vectorization result of the sample transaction associated data obtained by mapping the sample transaction associated data to a predetermined vector space.

[0209] In the specific implementation of this embodiment, for each sample transaction, the sample transaction-related data of the sample transaction is first mapped from the data space to a preset vector space, converting the sample transaction-related data into a data format acceptable to the model. Next, the representation of the sample transaction-related data in the vector space is used as the mapping result of the sample transaction-related data, and the sample transaction characteristics of the sample transaction are determined based on the mapping result of the sample transaction.

[0210] In step 830 , for each sample transaction, a sample reconstruction coefficient vector of the sample transaction is determined based on the sample transaction feature, the reference risk transaction feature, and the reference safety transaction feature.

[0211] To save space, the specific implementation process of determining the sample reconstruction coefficient vector of the sample transaction in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0212] In step 840 , for each sample transaction, the sample reconstruction coefficient vector is input into the logistic regression model, and a sample transaction risk score of the sample transaction is output.

[0213] To save space, the specific implementation process of determining the sample reconstruction coefficient vector of the sample transaction in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0214] In step 850 , based on the comparison of the sample transaction risk scores and the sample transaction labels of the plurality of sample transactions, the model parameters of the logistic regression model are adjusted to obtain a risk prediction model.

[0215] In a specific implementation of this embodiment, first, based on the sample transaction risk score of each sample transaction, a determination is made as to whether the sample transaction is predicted to be a risky transaction, thereby obtaining a predicted transaction label for the sample transaction. Next, based on the degree of difference between the sample transaction label and the predicted transaction label of the sample transaction, the model parameters of the logistic regression model are adjusted. Steps 810-850 are repeated to continuously reduce the difference between the sample transaction label and the predicted transaction label of the sample transaction until the difference between the sample transaction labels and the predicted transaction labels of multiple sample transactions meets the model training requirements. The model parameter update of the logistic regression model is then stopped, and the model parameters at this point are used as the final model parameters. The logistic regression model with the final model parameters is then used as the trained risk prediction model.

[0216] The benefit of this embodiment is that the logistic regression model is trained into a risk prediction model based on the LASSO algorithm, so that the risk prediction model has better nonlinearity, which can effectively improve the fitting ability and prediction ability of the risk prediction model. Compared with the method of using a neural network model in related technologies, the model structure of the risk prediction model of this method is simple. The prediction process does not require tedious calculations, does not rely on the model structure and calculation of the load, and is conducive to real-time deployment, so that when the risk prediction model is used for actual transaction risk prediction, it can take into account the accuracy and real-time requirements of transaction risk prediction.

[0217] The sample transaction associated data of the embodiment of the present disclosure includes the sample transaction content of the sample transaction and the sample transaction associated object.

[0218] The sample transaction content is used to indicate the specific transaction content of the sample transaction.

[0219] The sample transaction associated object is used to indicate the object associated with the sample transaction. The sample transaction associated data includes object information of the sample transaction associated object.

[0220] It should be noted that the object information of the sample transaction-related object includes but is not limited to the object location of the sample transaction-related object, the terminal address of the object, and the like.

[0221] Please refer to Figure 9 In one embodiment, step 820 specifically includes but is not limited to the following steps 910-930:

[0222] Step 910: Based on the mapping result of the sample transaction content, obtain the sample transaction content feature of the sample transaction;

[0223] Step 920: Based on the mapping result of the sample transaction associated object, obtain the sample transaction associated object feature of the sample transaction;

[0224] Step 930: Integrate the sample transaction content features and the sample transaction associated object features into sample transaction features.

[0225] Steps 910-930 are described in detail below.

[0226] In step 910 , based on the mapping result of the sample transaction content, the sample transaction content feature of the sample transaction is obtained.

[0227] The sample transaction content feature is a vectorized representation of the sample transaction content.

[0228] In a specific implementation of this embodiment, for the sample transaction content, the sample transaction content is mapped from the data space to a preset vector space to obtain a mapping result of the sample transaction content, and the mapping result is used as the sample transaction content feature of the sample transaction.

[0229] In step 920 , based on the mapping result of the sample transaction associated object, the sample transaction associated object feature of the sample transaction is obtained.

[0230] The sample transaction-related object feature is a vectorized representation of the basic object information of the sample transaction-related object.

[0231] In the specific implementation of this embodiment, for the sample transaction associated object, the object information of the sample transaction associated object is mapped from the data space to the preset vector space to obtain a mapping result of the sample transaction associated object, and the mapping result is used as the sample transaction associated object feature of the sample transaction.

[0232] In step 930 , the sample transaction content features and the sample transaction-associated object features are integrated into the sample transaction features.

[0233] In the specific implementation of this embodiment, the sample transaction content features and the sample transaction associated object features are fused to obtain a complete feature of a fixed length. This complete feature of the fixed length is used as the sample transaction feature so that the LASSO regression algorithm can be used to mine feature information of the complete feature of the fixed length, thereby obtaining a sample reconstruction coefficient vector of the sample transaction.

[0234] It should be noted that when mapping the sample transaction content and the sample transaction-associated objects to the vector space, the vector representation corresponding to each word in the sample transaction content and the vector representation corresponding to each word in the object information of the sample transaction-associated objects can be queried in a predetermined vocabulary. The vector representations corresponding to each word in the sample transaction content are integrated into the sample transaction content features of the sample transaction, and the vector representations corresponding to each word in the object information of the sample transaction-associated objects are integrated into the sample transaction-associated object features of the sample transaction. The predetermined vocabulary contains preset words and their corresponding vector representations. For example, the vector representation of the word "I" is "1." The vector representation of the word "delicious food" is "7,4."

[0235] The advantage of this embodiment is that the transaction information of the sample transaction, such as the sample transaction content and the sample transaction-related objects, is mapped to a preset vector space respectively, so as to obtain multiple types of mapping results for a sample transaction, and all the mapping results are subjected to feature fusion to obtain a sample transaction feature with more comprehensive feature information, which can better improve the feature quality of the sample transaction feature.

[0236] Please refer to Figure 10 In one embodiment, step 830 specifically includes but is not limited to the following steps 1010-1030:

[0237] Step 1010: Generate a predetermined transaction feature matrix based on the reference risk transaction features and the reference safety transaction features;

[0238] Step 1020: Determine an objective function based on a predetermined transaction feature matrix and a predetermined regularization coefficient;

[0239] Step 1030: For each sample transaction, substitute the sample transaction characteristics into the objective function, and determine the regularization term parameter that minimizes the output result of the objective function as the sample reconstruction coefficient vector of the sample transaction.

[0240] Steps 1010-1030 are described in detail below.

[0241] In step 1010, a predetermined transaction feature matrix is ​​generated based on reference risk transaction features and reference safety transaction features.

[0242] The predetermined transaction feature matrix is ​​used to store the reference risk transaction features and the reference safety transaction features in a vector form.

[0243] In a specific implementation of this embodiment, first, with authorization, a predetermined number of reference risk transactions are obtained from a preset channel, and reference risk transaction features are extracted for each reference risk transaction. Next, a predetermined number of reference safety transactions are obtained from a preset channel, and reference safety transaction features are extracted for each reference safety transaction, such that the number of features in the reference risk transaction features and the reference safety transaction features is equal. Furthermore, an all-zero matrix is ​​padded with the reference risk transaction features and the reference safety transaction features as elements to obtain a non-zero matrix, which is then determined as the predetermined transaction feature matrix.

[0244] For example, the predetermined number is set to N, where N is an integer greater than 0. The reference risk transaction feature of the j-th reference risk transaction obtained can be expressed as x 正,j The reference security transaction feature of the jth reference security transaction can be expressed as x 负,j Based on this, the scheduled transaction feature matrix can be expressed as:

[0245] A=[x 正,1 ,…,x 正,j ,…,x 正,N , x 负,1 ,…,x 负,j ,…,x 负,N ];

[0246] In step 1020 , an objective function is determined based on a predetermined transaction feature matrix and a predetermined regularization coefficient.

[0247] The objective function is a function based on the LASSO regression algorithm and the least square method, in which the LI regularization term is introduced. The predetermined regularization coefficient refers to the coefficient of the LI regularization term of the objective function.

[0248] In the specific implementation of this embodiment, first, the coefficients of the least squares term are determined based on the predetermined transaction feature matrix, and the coefficients of the LI regularization term of the objective function are determined based on the predetermined regularization coefficient. Then, the objective function is determined based on the function objective of the LASSO regression algorithm.

[0249] In step 1030, for each sample transaction, the sample transaction characteristics are substituted into the objective function, and the regularization term parameter that minimizes the output result of the objective function is determined as the sample reconstruction coefficient vector of the sample transaction.

[0250] In the specific implementation of this embodiment, the goal is to minimize the output of the objective function. By using the gradient descent method, the reconstruction coefficient that minimizes the output of the objective function is used as the sample reconstruction coefficient vector of the sample transaction. For each sample transaction, when the input sample transaction feature x iAfterwards, the reconstruction coefficient is solved according to the gradient descent method, and the reconstruction coefficient that minimizes the output result of the objective function is used as the sample reconstruction coefficient vector.

[0251] The calculation process of the sample reconstruction coefficient vector of the i-th sample transaction can be expressed as shown in formula (3):

[0252]

[0253] Among them, x i is the sample transaction feature of the i-th sample transaction, β i is the sample reconstruction coefficient vector of the i-th sample transaction, λ is the predetermined regularization coefficient, and A is the predetermined transaction feature matrix formed based on the reference risk transaction features and the reference security features.

[0254] The benefit of this embodiment is that, since in actual applications, the number of risky transactions is much smaller than the number of safe transactions, in order to improve the model's ability to learn and mine the transaction features of risky transactions, a predetermined transaction feature matrix is ​​set to influence the extraction process of sample transaction features of sample transactions. This can effectively reduce the impact of the difference in the proportion of safe transactions and risky transactions in the real world on feature selection and model training, and improve the vector quality of the sample reconstruction coefficient vector.

[0255] In one embodiment, the predetermined regularization coefficient is determined by:

[0256] Determine the sparsity of the sample reconstruction coefficient vector;

[0257] Based on the sparsity, a predetermined regularization coefficient is determined from among a plurality of candidate regularization coefficients.

[0258] In the specific implementation of this embodiment, λ is a hyperparameter that controls the regularization strength. When the predetermined regularization coefficient λ is larger, the sparsity of the sample reconstruction coefficient vector is larger; when the predetermined regularization coefficient λ is smaller, the sparsity of the sample reconstruction coefficient vector is smaller. Based on this, in order to improve the model training effect, the specific value of the predetermined regularization coefficient can be controlled according to the expected sparsity of the sample reconstruction coefficient vector of the sample transaction. First, according to the model training requirements, the sparsity of the sample reconstruction coefficient vector is determined. Then, the preset mapping table is called, wherein the preset mapping table is used to indicate the candidate regularization coefficients corresponding to each sparsity interval. Further, the sparsity interval in which the sparsity of the sample reconstruction coefficient vector is located is determined in the preset mapping table, and the candidate regularization coefficient corresponding to the sparsity interval in which the sparsity of the sample reconstruction coefficient vector is located is used as the predetermined regularization coefficient. Among them, the predetermined regularization coefficient of the embodiment of the present disclosure can take a value of 0.1.

[0259] In another embodiment, the predetermined regularization coefficient may be determined by function calculation. Specifically, first, a preset function for characterizing the relationship between sparsity and the regularization coefficient is called, the sparsity is substituted into the preset function, and the output of the preset function is used as the predetermined regularization coefficient, wherein the preset function is an increasing function with sparsity as the independent variable and the predetermined regularization coefficient as the dependent variable.

[0260] The advantage of this embodiment is that the shortcomings of the predetermined regularization coefficient are associated with the sparsity of the sample reconstruction coefficient vector, so that a suitable predetermined regularization coefficient can be determined based on the sparsity that the sample reconstruction coefficient vector must meet based on various methods such as table lookup or function calculation, thereby improving the accuracy of determining the predetermined regularization coefficient.

[0261] In the disclosed embodiment, the logistic regression model includes an activation function, a first parameter, and a second parameter. The activation function is used to calculate the probability of a sample transaction to determine whether the sample transaction is a risky transaction or a safe transaction. The activation function can be a sigmoid function, etc. The first parameter and the second parameter are model parameters in the logistic regression model and are trainable parameters. In each iterative training round, the first parameter and the second parameter can be updated so that the model prediction results ultimately output based on the first parameter and the second parameter meet the model training requirements. After the logistic regression model is trained into a risk prediction model, the first parameter and the second parameter are fixed.

[0262] Please refer to Figure 11 In one embodiment, step 840 specifically includes but is not limited to the following steps 1110-1120:

[0263] Step 1110: Input the sample reconstruction coefficient vector into the logistic regression model, and obtain a first calculation result based on the transposed result of the first parameter and the product of the sample reconstruction coefficient vector;

[0264] Step 1120: Probability score the sum of the first calculation result and the second parameter based on the activation function to obtain a sample transaction risk score of the sample transaction.

[0265] Steps 1110 - 1120 are described in detail below.

[0266] In step 1110, the sample reconstruction coefficient vector is input into the logistic regression model, and a first calculation result is obtained based on the transposed result of the first parameter and the product of the sample reconstruction coefficient vector.

[0267] In a specific implementation of this embodiment, first, the sample reconstruction coefficient vector is input into the logistic regression model, and a transposition operation is performed on the first parameter to obtain a transposed result of the first parameter. Then, the logistic regression model multiplies the sample reconstruction coefficient vector and the transposed result of the first parameter to obtain a first calculation result.

[0268] In step 1120 , a probability score is performed on the sum of the first calculation result and the second parameter based on the activation function to obtain a sample transaction risk score of the sample transaction.

[0269] In a specific implementation of this embodiment, the first calculation result and the second parameter are first added together to obtain the sum of the first calculation result and the second parameter. Next, the sum of the first calculation result and the second parameter is substituted into an activation function to perform activation calculation. The probability vector indicating that the sample transaction is a risky transaction in the probability distribution output by the activation function is used as the sample transaction risk score for the sample transaction.

[0270] The sample transaction risk score of the i-th sample transaction can be expressed as shown in formula (4):

[0271] p i =sigmoid(w T β i +b) Formula (4)

[0272] Among them, p i is the sample transaction risk score of the i-th sample transaction, β i is the sample reconstruction coefficient vector of the i-th sample transaction, w is the first parameter, w T is the transposed result of the first parameter. b is the second parameter.

[0273] The advantage of this embodiment is that the logistic regression model's calculation method is relatively simple, and it can calculate the sample transaction risk score of the sample transaction based on the activation function, the first parameter, and the second parameter at a relatively fast calculation speed, which is conducive to improving the real-time performance of transaction risk prediction. In addition, the logistic regression model receives the sample reconstruction coefficient vector obtained through feature selection, rather than the sample transaction features of the sample transaction, and does not directly input simple sample transaction features. This can effectively reduce the interference of invalid features on the transaction feature information learned by the logistic regression model, which is conducive to improving the fitting and predictive capabilities of the logistic regression model.

[0274] Because selecting different loss functions as sub-loss functions can lead to significant differences in the model's training results, the disclosed embodiments provide a calculation scheme based on the logarithmic loss function, which can improve the computational efficiency of the total loss function and the accuracy of measuring the model's training results using the total loss function.

[0275] Please refer to Figure 12In one embodiment, step 850 specifically includes but is not limited to the following steps 1210-1230:

[0276] Step 1210: For each sample transaction, calculate a sub-loss function based on the sample transaction risk score and the sample transaction label;

[0277] Step 1220: Obtain a total loss function based on the sum of multiple sub-loss functions;

[0278] Step 1230: Adjust the model parameters of the logistic regression model based on the total loss function to obtain a risk prediction model.

[0279] Steps 1210 - 1230 are described in detail below.

[0280] In step 1210 , for each sample transaction, a sub-loss function is calculated based on the sample transaction risk score and the sample transaction label.

[0281] The sample transactions in the embodiment of the present disclosure include positive sample transactions and negative sample transactions.

[0282] A positive sample transaction refers to a sample transaction whose sample transaction label is 1 among multiple sample transactions, that is, a positive sample transaction is a risky transaction among multiple sample transactions.

[0283] A negative sample transaction refers to a sample transaction whose sample transaction label is 0 among multiple sample transactions, that is, a negative sample transaction is a safe transaction among multiple sample transactions.

[0284] In a specific implementation of this embodiment, first, based on the sample transaction's sample transaction label, a determination is made as to whether the sample transaction is a positive or negative sample transaction. Next, if the sample transaction is determined to be a positive sample transaction, a sub-loss function is determined based on the negative logarithm of the sample transaction's risk score. If the sample transaction is determined to be a negative sample transaction, a sub-loss function is determined based on the negative logarithm of the difference between 1 and the sample transaction's risk score.

[0285] In step 1220, a total loss function is obtained based on the sum of multiple sub-loss functions.

[0286] The total loss function is used to measure the prediction accuracy of the risk prediction model. The smaller the total loss function, the higher the transaction risk prediction accuracy of the risk prediction model.

[0287] In the specific implementation of this embodiment, since the sub-loss function indicates the difference between the real risk label and the predicted risk label of each sample transaction, based on this, first, the total number of transactions of the sample transactions is determined. Then, based on the total number of transactions, the sub-loss functions of all sample transactions are added together to obtain the sum of multiple sub-loss functions, and the sum of multiple sub-loss functions is determined as the total loss function. Among them, the total loss function of the embodiment of the present disclosure can be expressed as shown in formula (5):

[0288]

[0289] Where Loss refers to the total loss function, n refers to the total number of sample transactions, and n is an integer greater than 0. i refers to the i-th sample transaction, and i is an integer in the range of (0, n]. i Refers to the sample transaction label of the i-th sample transaction, y i The value of p is 0 or 1. i Refers to the sample transaction risk score of the i-th sample transaction, p i The value range is [0,1].

[0290] In step 1230, the model parameters of the logistic regression model are adjusted based on the total loss function to obtain a risk prediction model.

[0291] In the specific implementation of this embodiment, the model parameters of the logistic regression model are adjusted with minimizing the total loss function as the training goal, and the above steps 810-850 are repeated to achieve iterative training of the logistic regression model. The model parameters that minimize the total loss function are used as the final model parameters, and the logistic regression model with the final model parameters is used as the trained risk prediction model.

[0292] The benefit of this embodiment is that by setting the sample transaction labels to 0 or 1, the sample transactions are divided into positive sample transactions and negative sample transactions, thereby converting the calculation process of the total loss function into the calculation process of the binary classification loss function, which can provide more accurate classification prediction results for model training, and at the same time reduce some unnecessary calculations brought about by using the cross entropy loss function as the total loss function, thereby improving calculation efficiency and accuracy.

[0293] like Figure 13As shown, it is an overall flow chart of the model training of the embodiment of the present disclosure. Specifically, first, a plurality of reference risk transactions and a plurality of reference security transactions are obtained, and a dictionary sample set is formed based on the plurality of reference risk transactions and reference security transactions, and the reference risk transaction features and reference security transaction features contained in the dictionary sample set are stored into a matrix to obtain a predetermined transaction feature matrix, and its specific implementation process is similar to the above step 1010. Next, a plurality of sample transactions are obtained, and a training sample set is formed based on the plurality of sample transactions, and its specific implementation process is similar to the above step 810. Further, the sample reconstruction coefficient vector (LASSO reconstruction coefficient vector) of each sample transaction is calculated, and its specific implementation process is similar to the above steps 1020-1030. Next, the sample reconstruction coefficient vector of each sample transaction is used to train the logistic regression model to obtain a risk prediction model, so that the trained risk prediction model can be used to detect risk transactions, and its specific implementation process is similar to the above steps 840-850.

[0294] Detailed description of model evaluation of a risk prediction model according to an embodiment of the present disclosure

[0295] In order to make the risk prediction model used for transaction risk prediction have better stability, the embodiment of the present disclosure provides a solution for evaluating the trained risk prediction model based on test transactions. When the model performance of the risk prediction model is relatively stable, the risk prediction model can be deployed online to improve the stability of the model application.

[0296] Please refer to Figure 14 In one embodiment, the process of evaluating the risk prediction model specifically includes but is not limited to the following steps 1410-1430:

[0297] Step 1410: Obtain first transaction features and first transaction tags of multiple first test transactions;

[0298] Step 1420: For each first test transaction, input the first transaction feature of the first test transaction into the risk prediction model to obtain a first transaction risk score for the first test transaction.

[0299] Step 1430: Based on the first transaction risk scores and first transaction labels of the plurality of first test transactions, perform a first model verification on the risk prediction model, so as to update the model parameters of the risk prediction model when it is determined that the first model verification fails.

[0300] Steps 1410-1430 are described in detail below.

[0301] In step 1410, first transaction features and first transaction tags of a plurality of first test transactions are obtained.

[0302] The first test transaction refers to a transaction used to perform a first model verification on the risk prediction model.

[0303] The first transaction feature is used to indicate the specific transaction content of the first test transaction, object feature information of the transaction-related object, etc.

[0304] The first transaction tag is used to identify whether the first test transaction is actually a risky transaction. When the first test transaction is a risky transaction, the first transaction tag is 1; when the first test transaction is a safe transaction (non-risky transaction), the first transaction tag is 0.

[0305] In the specific implementation of this embodiment, the specific implementation process of step 1410 is similar to the specific implementation process of step 810 described above. The difference is that the sample transaction in step 810 is used for model training, while the first test transaction in step 1410 is used for model evaluation. The two have different purposes. To save space, they are not further described.

[0306] In step 1420, for each first test transaction, the first transaction feature of the first test transaction is input into the risk prediction model to obtain a first transaction risk score of the first test transaction.

[0307] In the specific implementation of this embodiment, the specific implementation process of step 1420 is similar to the specific implementation process of step 320 and steps 830-840 described above. The difference is that step 320 is used in the application phase of the risk prediction model, steps 830-840 are used in the training phase of the risk prediction model, and step 1420 is used in the evaluation and testing phase of the risk prediction model. To save space, they are not repeated here.

[0308] In step 1430, based on the first transaction risk scores and first transaction labels of the plurality of first test transactions, a first model verification is performed on the risk prediction model, so that when it is determined that the first model verification fails, the model parameters of the risk prediction model are updated.

[0309] When this embodiment is specifically implemented, first, the predicted transaction label of the first test transaction is determined based on the first transaction risk score of the first test transaction. Then, based on the predicted transaction labels of multiple first test transactions and the difference between the first transaction labels, the model indicators such as accuracy and precision of the risk prediction model are calculated. Furthermore, the first model is verified based on the calculated model indicators and the conditions to be met by the model indicators of the risk prediction model. When the model indicators of the risk prediction model all meet the corresponding conditions, it is determined that the first model verification has passed. When there is a model indicator of the risk prediction model that does not meet the corresponding conditions, it is determined that the first model verification has failed, and returns to steps 810-850 to update the model parameters of the risk prediction model to continue training the risk prediction model.

[0310] The benefit of this embodiment is that a model evaluation is performed on the risk prediction model obtained through training based on the first test transaction, and the model stability and model accuracy of the risk prediction model are evaluated, so that when the model performance of the risk prediction model is relatively stable (when the first model verification is passed), the risk prediction model can be deployed online to improve the stability of the model application, and when the first model verification fails, the model parameters of the risk prediction model are updated to further improve the model performance of the risk prediction model.

[0311] In order to improve the efficiency of model evaluation, it is usually adopted that when obtaining sample transactions as training samples, the obtained training sample data is divided into a training set and a test set in proportion, the training samples of the training set are used as the sample transactions of the above steps 810-850 for model training, and the training samples of the test set are used as the first test transactions for the first model verification. This method often leads to the difference between the training samples based on model training and model evaluation. It is often difficult to evaluate the transaction risk prediction ability of the risk prediction model in the time dimension only through the first model verification. Based on this, the embodiment of the present disclosure provides a solution for performing performance and stability evaluation on the risk prediction model that has passed the first model verification based on cross-time samples, which can further evaluate the model performance and stability of the risk prediction model.

[0312] Please refer to Figure 15 In one embodiment, after the first model is verified, the transaction risk prediction method further includes but is not limited to the following steps 1510-1530:

[0313] Step 1510: If it is determined that the first model verification passes, obtain second transaction features and second transaction tags of multiple second test transactions;

[0314] Step 1520: For each second test transaction, input the second transaction feature of the second test transaction into the risk prediction model to obtain a second transaction risk score for the second test transaction.

[0315] Step 1530: Based on the second transaction risk scores and second transaction labels of multiple second test transactions, perform a second model verification on the risk prediction model, so as to update the model parameters of the risk prediction model when it is determined that the second model verification fails, or deploy the risk prediction model online when the second model verification passes.

[0316] Steps 1510-1530 are described in detail below.

[0317] In step 1510 , if it is determined that the first model verification passes, second transaction features and second transaction tags of a plurality of second test transactions are obtained.

[0318] The second test transaction refers to a transaction used to perform a second model verification on the risk prediction model.

[0319] The second transaction feature is used to indicate the specific transaction content of the second test transaction, object feature information of the transaction-related object, etc.

[0320] The second transaction tag is used to identify whether the second test transaction is actually a risky transaction. When the second test transaction is a risky transaction, the second transaction tag is 1; when the second test transaction is a safe transaction (non-risky transaction), the second transaction tag is 0.

[0321] The transaction time difference between the second test transaction and the first test transaction in the embodiment of the present disclosure meets a predetermined condition.

[0322] The predetermined condition is used to define an extreme value of a transaction time difference between the second test transaction and the first test transaction.

[0323] During the specific implementation of this embodiment, if it is determined that the first model verification has passed, it indicates that the transaction risk prediction capability of the trained risk prediction model on the first test transaction meets the predetermined requirements. Furthermore, in order to improve the evaluation effect of the risk prediction model, it is also necessary to evaluate the transaction risk prediction capability of the risk prediction model on cross-time sample transactions. Based on this, with authorization, the second transaction features and second transaction labels of multiple second test transactions are obtained, and the specific implementation process is similar to the specific implementation process of step 1410 above. The difference is that there is a large difference in the time dimension between the first test transaction and the second test transaction. To save space, it will not be described in detail.

[0324] In step 1520, for each second test transaction, the second transaction feature of the second test transaction is input into the risk prediction model to obtain a second transaction risk score of the second test transaction.

[0325] In the specific implementation of this embodiment, the specific implementation process of step 1520 is similar to the specific implementation process of step 1420 described above. The difference is that step 1420 predicts the risk of the first test transaction during the model evaluation phase, while step 1520 predicts the risk of the second test transaction during the model evaluation phase. The two processes use different test samples. To save space, they are not further described.

[0326] In step 1530, based on the second transaction risk scores and second transaction labels of multiple second test transactions, the risk prediction model is subjected to a second model verification, so as to update the model parameters of the risk prediction model when it is determined that the second model verification fails, or to deploy the risk prediction model online when the second model verification passes.

[0327] In the specific implementation of this embodiment, the specific implementation process of step 1530 is similar to the specific implementation process of the above-mentioned step 1430. To save space, it is not repeated here.

[0328] The benefit of this embodiment is that the performance evaluation and stability evaluation of the risk prediction model that has passed the first model verification are performed based on cross-time samples (multiple second test transactions), which can further evaluate the model performance and stability of the risk prediction model, and is conducive to improving the comprehensiveness and accuracy of the model evaluation.

[0329] Please refer to Figure 16 In one embodiment, the predetermined condition satisfied by the transaction time difference between the second test transaction and the first test transaction is determined by:

[0330] Step 1610: Determine the training accuracy of the risk prediction model and the total number of predetermined transactions for the second test transaction;

[0331] Step 1620: Based on the training accuracy and the total number of predetermined transactions, determine the extreme value of the transaction time difference between the second test transaction and the first test transaction, and determine the predetermined condition based on the extreme value of the time difference.

[0332] Steps 1610-1620 are described in detail below.

[0333] In step 1610, first, with authorization, multiple set training parameters for the risk prediction model are obtained. Then, the training accuracy of the risk prediction model and the total number of predetermined transactions for the second test transaction are screened out from the multiple set training parameters.

[0334] In step 1620, first, in a mapping table used to characterize the correspondence between training accuracy and transaction time difference intervals, the transaction time difference interval corresponding to the training accuracy of the risk prediction model is searched as the first time difference interval. Next, in a mapping table used to characterize the correspondence between the total number of predetermined transactions and transaction time difference intervals, the transaction time difference interval corresponding to the total number of predetermined transactions of the second test transaction is searched as the second time difference interval. Furthermore, the intersection of the first time difference interval and the second time difference interval is taken to obtain a target time difference interval, and based on the time endpoints of the target time difference interval, the transaction time difference extreme value between the second test transaction and the first test transaction is determined, wherein the transaction time difference extreme value includes a maximum time difference and a minimum time difference. Finally, the time interval defined by the transaction time difference extreme value is determined as the predetermined condition that the transaction time difference between the second test transaction and the first test transaction must meet.

[0335] The advantage of this embodiment is that the training accuracy of the combined risk prediction model and the total number of second test transactions that must be met are combined, and score calculation and table lookup are introduced to jointly determine the extreme value of the transaction time difference between the second test transaction and the first test transaction. This can clearly and quickly define the selection range of the second test transaction in the time dimension, which is conducive to improving the efficiency of obtaining the second test transaction.

[0336] Please refer to Figure 17 In one embodiment, step 1530 specifically includes but is not limited to the following steps 1710-1730:

[0337] Step 1710: Determine a confusion matrix of the risk prediction model based on the second transaction risk scores, second transaction labels, and a predetermined threshold of the plurality of second test transactions;

[0338] Step 1720: Perform a first sub-validation on the risk prediction model based on the confusion matrix according to the first validation rule.

[0339] Step 1730: According to the second verification rule and based on the confusion matrix, perform a second sub-verification on the risk prediction model.

[0340] Steps 1710-1730 are described in detail below.

[0341] In step 1710, a confusion matrix of the risk prediction model is determined based on the second transaction risk scores of the plurality of second test transactions, the second transaction labels, and a predetermined threshold.

[0342] The predetermined threshold is used to predict the second test transaction as a risky transaction or a safe transaction in combination with the second transaction risk score. The predetermined threshold has a value range of [0, 1].

[0343] The confusion matrix is ​​used to display the classification results of the risk prediction model for each second test transaction in matrix form.

[0344] In the specific implementation of this embodiment, the confusion matrix of the disclosed embodiment is a 2×2 matrix comprising four matrix elements. Specifically, first, for each second test transaction, the second transaction risk score is compared with a predetermined threshold. Next, the predicted transaction type of the second test transaction whose second transaction risk score is greater than or equal to the predetermined threshold is determined to be a risky transaction; the predicted transaction type of the second test transaction whose second transaction risk score is less than the predetermined threshold is determined to be a safe transaction. Furthermore, based on the predicted transaction type and second transaction label of the second test transaction, the total number of second test transactions whose predicted transaction type and second transaction label are both risky transactions is determined as the first element TP (True Positive) in the confusion matrix; the total number of second test transactions whose predicted transaction type and second transaction label are both safe transactions is determined as the second element TN (True Negative) in the confusion matrix; the total number of second test transactions whose predicted transaction type is a risky transaction but whose second transaction label is a safe transaction is determined as the third element FP (False Positive) in the confusion matrix; and the total number of second test transactions whose predicted transaction type is a safe transaction but whose second transaction label is a risky transaction is determined as the fourth element FN (False Negative) in the confusion matrix. Finally, a confusion matrix of the risk prediction model is generated based on the first element, the second element, the third element, and the fourth element.

[0345] In step 1720, according to the first verification rule, a first sub-verification is performed on the risk prediction model based on the confusion matrix.

[0346] The first verification rule is used to limit the test indicators that need to be clarified in the first sub-verification of the risk prediction model, the calculation method of each test indicator, and the requirements that each test indicator must meet.

[0347] To save space, the specific process of performing the first sub-verification of the risk prediction model in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0348] In step 1730, according to the second verification rule, a second sub-verification is performed on the risk prediction model based on the confusion matrix.

[0349] The second verification rule is used to limit the test indicators that the risk prediction model needs to clarify in the second sub-verification, the calculation method of each test indicator, and the requirements that each test indicator must meet.

[0350] To save space, the specific process of performing the second sub-verification of the risk prediction model in the embodiment of the present disclosure will be described in detail below and will not be repeated here.

[0351] The benefit of this embodiment is that it introduces the concept of confusion matrix to perform model evaluation on the risk prediction model, and performs multiple verifications on the risk prediction model based on the characteristic elements in the confusion matrix and multiple verification rules, which can improve the comprehensiveness and accuracy of the model evaluation of the risk prediction model.

[0352] In one embodiment, the predetermined threshold value of step 1710 is determined by:

[0353] Determine the current application scenario type of the risk prediction model;

[0354] Based on the current application scenario type, a predetermined threshold is determined from a plurality of candidate thresholds.

[0355] In the specific implementation of this embodiment, first, with authorization, multiple set training parameters for the risk prediction model are obtained. Then, the current application scenario type of the risk prediction model is screened out from the multiple set training parameters. Finally, in a mapping table used to characterize the correspondence between model application scenarios and candidate thresholds, the model application scenario corresponding to the current application scenario type is searched, and the candidate threshold corresponding to the model application scenario consistent with the current application scenario type is used as the predetermined threshold.

[0356] The advantage of this embodiment is that, based on the scenario type to which the risk prediction model is applied, the predetermined threshold is set to a value that better matches the application scenario, which can meet the transaction risk prediction requirements of different application scenarios and improve the applicability of the risk prediction model.

[0357] Please refer to Figure 18 In some embodiments, step 1720 specifically includes but is not limited to the following steps 1810-1820:

[0358] Step 1810: Determine the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy of the risk prediction model based on the confusion matrix;

[0359] Step 1820: Perform a first sub-verification on the risk prediction model based on the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy.

[0360] Steps 1810-1820 are described in detail below.

[0361] In step 1810, based on the confusion matrix, the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy of the risk prediction model are determined.

[0362] The information ratio is used to indicate the degree of difference between the second transaction risk scores of risky transactions and safe transactions predicted by the risk prediction model.

[0363] The maximum cumulative transaction risk score difference is used to indicate the upper limit of the second transaction risk scores of risky transactions and safe transactions predicted by the risk prediction model.

[0364] Feature discrimination accuracy is used to indicate the degree to which the risk prediction model distinguishes between risky and safe transactions.

[0365] In the specific implementation of this embodiment, when calculating the information ratio of the risk prediction model, first, the total number of the second test transactions is determined, then the first proportion of the second test transactions predicted by the risk prediction model as risky transactions and with the second transaction label as risky transactions in the second test transactions with the second transaction label as risky transactions is determined, and the second proportion of the second test transactions predicted by the risk prediction model as safe transactions and with the second transaction label as safe transactions in the second test transactions with the second transaction label as safe transactions is determined. Finally, the information ratio of the risk prediction model is obtained based on the difference between the first proportion and the second proportion and the product of the logarithm of the quotient of the first proportion and the second proportion.

[0366] Furthermore, when calculating the maximum cumulative transaction risk score difference, the difference between the first proportion and the second proportion is taken as the maximum cumulative transaction risk score difference.

[0367] Furthermore, when calculating feature discrimination accuracy, first, based on the second transaction score of each second test transaction, the distribution function of the risk transaction predicted by the risk prediction model and the distribution function of the safety transaction predicted by the risk prediction model are determined, where both of the aforementioned distribution functions are cumulative probability density functions. Next, the product of the distribution function of the predicted risk transaction and the distribution function of the predicted safety transaction is integrated between positive infinity and negative infinity to obtain the feature discrimination accuracy.

[0368] In step 1820, a first sub-validation is performed on the risk prediction model based on the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy.

[0369] In a specific implementation of this embodiment, first, based on the first verification rule, the preset conditions that must be met by the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy are determined. Next, the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy actually calculated by the risk prediction model are compared with the preset conditions. When the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy all meet the preset conditions, the first sub-verification is determined to have passed. When at least one of the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy does not meet the preset conditions, the first sub-verification is determined to have failed.

[0370] The benefit of this embodiment is that, based on the first verification rule, the risk prediction model is evaluated by combining the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy, thereby achieving evaluation of the model in multiple dimensions such as accuracy and stability, and improving the comprehensiveness of the model evaluation.

[0371] Please refer to Figure 19 In one embodiment, step 1730 specifically includes but is not limited to the following steps 1910-1930:

[0372] Step 1910: Based on the confusion matrix, determine a first number of second test transactions whose second transaction risk scores are greater than or equal to a predetermined threshold, a second number of second test transactions whose second transaction label is 1, and a third number of second test transactions whose second transaction risk scores are greater than or equal to the predetermined threshold and whose second transaction label is 1.

[0373] Step 1920: According to the second verification rule, determine the transaction coverage of the risk prediction model based on the quotient of the third number and the second number, and determine the policy cost performance of the risk prediction model based on the quotient of the first number and the third number;

[0374] Step 1930: Perform a second sub-verification on the risk prediction model based on the transaction coverage and the strategy cost-effectiveness.

[0375] Steps 1910-1930 are described in detail below.

[0376] The transaction coverage rate is used to indicate the proportion of real risk transactions covered by the risk prediction model in all real risk transactions.

[0377] The strategy cost-effectiveness is used to indicate the multiple of the risk events covered by the risk prediction model and the actual risk events covered by the risk prediction model.

[0378] In step 1910, based on the elements of the confusion matrix, first, determine the first number of second test transactions whose second transaction risk scores are greater than or equal to a predetermined threshold. The first number is the total number of risk transactions affected by the risk prediction model, where the first number is the sum of the third element FP and the first element TP. Next, determine the second number of second test transactions whose second transaction label is 1. The second number is the total number of second test transactions that are truly risky transactions, where the second number is the sum of the first element TP and the fourth element FN. Finally, determine the third number of second test transactions whose second transaction risk scores are greater than or equal to the predetermined threshold and whose second transaction label is 1. The third number is the total number of risk transactions covered by the risk prediction model, where the third number is the first element TP.

[0379] In step 1920, first, according to the second verification rule, the transaction coverage of the risk prediction model is determined based on the quotient of the third number and the second number, where transaction coverage = (first element TP) / (first element TP + fourth element FN). Then, according to the second verification rule, the policy cost performance of the risk prediction model is determined based on the quotient of the first number and the third number, where policy cost performance = (third element FP + first element TP) / (first element TP).

[0380] In step 1930, based on the second verification rule, the pre-set conditions that transaction coverage and policy cost-effectiveness must meet are determined. Next, the transaction coverage and policy cost-effectiveness actually calculated by the risk prediction model are compared with the pre-set conditions. If both the transaction coverage and policy cost-effectiveness meet the pre-set conditions, the second sub-verification is determined to have passed. If at least one of the transaction coverage and policy cost-effectiveness does not meet the pre-set conditions, the second sub-verification is determined to have failed.

[0381] The benefit of this embodiment is that, based on the second verification rule, the proportion of actual risk affairs covered by the risk prediction model in all actual risk affairs is determined, and the risk affairs covered by the risk prediction model and the multiples of the actual risk affairs covered by the risk prediction model are determined. The risk prediction model is evaluated together with the transaction coverage rate and the strategy cost-effectiveness, which can improve the comprehensiveness of the model evaluation.

[0382] The data collection process of the model training and model application stages in the embodiment of the present disclosure

[0383] In real-world scenarios, model training and application are significantly impacted by environmental factors, including the current network status. To improve the timeliness of updating risk prediction models, this disclosure provides a method for collecting log data from the model training and application phases, so that the risk prediction model can be updated when the log data meets predetermined conditions.

[0384] Please refer to Figure 20 In one embodiment, the data collection process for the model training and model application phases specifically includes but is not limited to the following steps 2010-2020:

[0385] Step 2010: Determine data collection conditions for the risk prediction model;

[0386] Step 2020: If it is determined that the model training and model operation of the risk prediction model meet the data collection conditions, the log data of the risk prediction model is obtained to update the model based on the log data.

[0387] Steps 2010-2020 are described in detail below.

[0388] In step 2010, data collection conditions for the risk prediction model are determined.

[0389] Data collection conditions are used to limit the specific time when data collection is required during the data collection model training and model application stages.

[0390] In the specific implementation of this embodiment, since the prediction accuracy of the risk prediction model depends to a large extent on the training samples and model training. Based on this, first of all, the acquisition of training samples and the training of the logistic regression model in the model training phase are used as data collection conditions for the risk prediction model, so that data collection is performed when the training samples are acquired and the logistic regression model is trained. Furthermore, after the model is deployed online, the operation of the risk prediction model also needs to be paid attention to. Based on this, the operation of the model after it is put online is also used as a data collection condition for the risk prediction model, so that the log data of the model during operation is collected in real time.

[0391] In step 2020, if it is determined that the model training and model operation of the risk prediction model meet the data collection conditions, the log data of the risk prediction model is obtained to update the model based on the log data.

[0392] Log data refers to the data generated by the risk prediction model during the model training phase and the model operation phase.

[0393] In a specific implementation of this embodiment, if it is determined that the model training and model operation of the risk prediction model meet the data collection conditions, then, with authorization, log data generated by the risk prediction model during model training and log data generated by the risk prediction model during model operation are collected. Subsequently, the collected log data is used to generate a report according to a predetermined report template to form report data specific to the risk prediction model. The model training and operation status are analyzed based on the report data, so that the risk prediction model can be updated if anomalies occur in the model training or operation.

[0394] like Figure 21As shown, for the risk prediction model, the steps of dictionary sample acquisition, training sample acquisition, LASSO reconstruction calculation, logistic regression model construction, and model training are first performed in sequence. The specific implementation process is similar to steps 810-850 described above. Next, the trained risk prediction model is evaluated. If the model evaluation fails, the process returns to the dictionary sample acquisition step. After the model evaluation, the model enters online development. The specific implementation process is similar to steps 1410-1430 and 1510-1530 described above. Furthermore, after the online model development is completed, the model is launched online, and risk prediction is performed online for each target transaction based on the risk prediction model. The specific implementation process is similar to steps 310-320 described above. In addition, during the dictionary sample acquisition, model training, and model launch phases, data collection and reporting steps are added to facilitate timely model updates. The specific implementation process is similar to steps 2010-2020 described above. To save space, this will not be further described.

[0395] The benefit of this embodiment is that, based on preset data collection conditions, when the data collection conditions are met, certain log data from the model training and model application stages is collected, and the model training and operation status are analyzed based on the collected log data. This allows data collection to be performed at more critical time points, reducing the complexity of data collection and the amount of collected log data. Furthermore, by analyzing the model training and operation status based on the amount of log data at more critical time points, the risk prediction model can be updated and repaired in a timely manner when anomalies occur in the model, thereby improving the stability of model training and model operation.

[0396] Implementation details of a transaction risk prediction method according to an embodiment of the present disclosure

[0397] Refer to the following Figures 22A-22B , which illustrates in detail the implementation details of the transaction risk prediction method of the embodiment of the present disclosure.

[0398] like Figure 22AFigure 1 illustrates the interaction between the transaction processing platform server 110 and the target terminal 140. During the model training phase, the transaction processing platform server 110 first performs LASSO reconstruction calculations based on reference shared transactions, reference secure transactions, and sample transactions. The specific implementation process is similar to the specific implementation process of steps 1010-1030 described above. Next, the sample reconstruction coefficient vector is input into the logistic regression model. The logistic regression model outputs a transaction risk score for the sample transaction based on the sample reconstruction coefficient vector. The specific implementation process is similar to the specific implementation process of steps 1110-1120 described above. Furthermore, a first loss function is generated based on the transaction risk score of the sample transaction from the first category prediction result and the sample transaction label. The logistic regression model is then adjusted based on the first loss function to generate a risk prediction model. The specific implementation process is similar to the specific implementation process of steps 1210-1230 described above. Furthermore, when the target terminal 140 wishes to perform risk prediction on a target transaction, it sends the target transaction features of the target transaction to the transaction processing platform server 110. Next, the transaction processing platform server 110 inputs the received target transaction features into the risk prediction model to obtain the target transaction risk score. The specific implementation process is similar to the specific implementation process of step 320 above. Finally, the transaction processing platform server 110 feeds back the generated transaction processing results based on the target policy to the target terminal 140.

[0399] like Figure 22B Figure 2 shows a detailed schematic diagram of the training, deployment, and application of a risk prediction model. During the offline deployment phase, the dictionary and training samples are first acquired, with the specific implementation process being similar to steps 1010 and 810 described above. Next, the LASSO reconstruction coefficient is calculated, with the specific implementation process being similar to steps 1020-1030 described above. Furthermore, the logistic regression model is trained, with the specific implementation process being similar to steps 840-850, steps 1410-1450, and steps 1510-1530 described above. Finally, the generated risk prediction model is deployed online to the CKV+TSSD database, a high-performance database compatible with multiple protocols such as Redis and ASN. Furthermore, after the model is deployed online, the risk prediction model and risk control policy are matched at the risk control policy layer. The risk control policy and risk prediction model are then used together to identify risks in transactions submitted by objects on the transaction processing platform, with the specific implementation process being similar to steps 310-320 described above. To save space, I will not go into details.

[0400] Description of the apparatus and device of the present disclosure

[0401] It is to be understood that, although the steps in the above-mentioned flowcharts are shown in sequence according to the arrow representations, these steps are not necessarily performed in sequence according to the order represented by the arrows. Unless otherwise specified in the present embodiment, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flowcharts may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0402] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object, such as the target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of such data will comply with relevant laws, regulations and standards. In addition, when the embodiment of the present application needs to obtain the attribute information of the target object, the target object's separate permission or separate consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the target object's separate permission or separate consent, the necessary target object-related data for the normal operation of the embodiment of the present application will be obtained.

[0403] Figure 23 This is a schematic diagram of the structure of the transaction risk prediction device 2300 provided in an embodiment of the present disclosure. The transaction risk prediction device 2300 includes:

[0404] An acquiring unit 2310 is configured to acquire target transaction characteristics of a target transaction;

[0405] The prediction unit 2320 is configured to input the target transaction characteristics of the target transaction into the risk prediction model to obtain a target transaction risk score of the target transaction, and perform transaction processing on the target transaction based on the target transaction risk score;

[0406] Among them, the risk prediction model is obtained by training a preset logistic regression model based on the comparison of the sample transaction risk scores of multiple sample transactions and the sample transaction labels of the sample transactions; the sample transaction risk score is obtained by inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model; the sample reconstruction coefficient vector is predicted based on the sample transaction association data of the sample transaction, as well as the pre-set reference risk transaction characteristics and reference security transaction characteristics.

[0407] Optionally, the risk prediction model includes a reconstruction function and a prediction network;

[0408] The prediction unit 2320 is specifically configured to:

[0409] Calculate the target reconstruction coefficient vector of the target transaction based on the reference risk transaction characteristics, the reference safety transaction characteristics, and the target transaction characteristics through the reconstruction function;

[0410] The target reconstruction coefficient vector is input into the prediction network for probability calculation to obtain the target transaction risk score.

[0411] Optionally, performing transaction processing on the target transaction based on the target transaction risk score includes:

[0412] Invoke multiple predetermined candidate strategies and trigger conditions corresponding to each predetermined candidate strategy;

[0413] If it is determined that the target transaction risk score and the target transaction characteristics meet the triggering condition of a target policy among a plurality of predetermined candidate policies, transaction processing is performed on the target transaction based on the target policy.

[0414] Optionally, the transaction risk prediction device 2300 further includes a training unit (not shown), which specifically includes:

[0415] a sample acquisition unit (not shown), configured to acquire sample transaction association data and sample transaction tags of a plurality of sample transactions;

[0416] a mapping unit (not shown), configured to obtain, for each sample transaction, sample transaction features of the sample transaction based on a mapping result of the sample transaction-associated data;

[0417] a determination unit (not shown), configured to determine, for each sample transaction, a sample reconstruction coefficient vector of the sample transaction based on the sample transaction feature, the reference risk transaction feature, and the reference safety transaction feature;

[0418] An input unit (not shown) is used to input the sample reconstruction coefficient vector into the logistic regression model for each sample transaction, and output a sample transaction risk score for the sample transaction;

[0419] An adjustment unit (not shown) is used to adjust the model parameters of the logistic regression model based on the comparison of the sample transaction risk scores and sample transaction labels of multiple sample transactions to obtain a risk prediction model.

[0420] Optionally, the sample transaction associated data includes sample transaction content of the sample transaction and a sample transaction associated object;

[0421] The mapping unit (not shown) is used to:

[0422] Based on the mapping result of the sample transaction content, a sample transaction content feature of the sample transaction is obtained;

[0423] Based on the mapping result of the sample transaction associated object, a sample transaction associated object feature of the sample transaction is obtained;

[0424] The sample transaction content features and the sample transaction associated object features are integrated into the sample transaction features.

[0425] Optionally, the determining unit (not shown) is configured to:

[0426] Generate a predetermined transaction feature matrix based on reference risk transaction features and reference safety transaction features;

[0427] Determining an objective function based on a predetermined transaction feature matrix and a predetermined regularization coefficient;

[0428] For each sample transaction, the sample transaction characteristics are substituted into the objective function, and the regularization term parameter that minimizes the output result of the objective function is determined as the sample reconstruction coefficient vector of the sample transaction.

[0429] Optionally, the logistic regression model includes an activation function, a first parameter, and a second parameter;

[0430] The input unit (not shown) is used to:

[0431] Inputting the sample reconstruction coefficient vector into the logistic regression model, and obtaining a first calculation result based on the transposed result of the first parameter and the product of the sample reconstruction coefficient vector;

[0432] A probability score is performed on the sum of the first calculation result and the second parameter based on the activation function to obtain a sample transaction risk score of the sample transaction.

[0433] Optionally, the adjustment unit (not shown) is configured to:

[0434] For each sample transaction, a sub-loss function is calculated based on the sample transaction risk score and sample transaction label;

[0435] Based on the sum of multiple sub-loss functions, the total loss function is obtained;

[0436] The model parameters of the logistic regression model are adjusted based on the total loss function to obtain the risk prediction model.

[0437] Optionally, the sample transactions include positive sample transactions and negative sample transactions;

[0438] For each sample transaction, based on the sample transaction risk score and sample transaction label, a sub-loss function is calculated, including:

[0439] If the sample transaction is determined to be a positive sample transaction, a sub-loss function is determined based on the negative logarithm of the sample transaction risk score;

[0440] If the sample transaction is determined to be a negative sample transaction, a sub-loss function is determined based on the negative logarithm of the difference between 1 and the risk score of the sample transaction.

[0441] Optionally, the transaction risk prediction device 2300 further includes a first testing unit (not shown), which is specifically configured to:

[0442] Obtaining first transaction features and first transaction tags of a plurality of first test transactions;

[0443] For each first test transaction, inputting the first transaction feature of the first test transaction into the risk prediction model to obtain a first transaction risk score of the first test transaction;

[0444] Based on the first transaction risk scores and first transaction labels of the plurality of first test transactions, a first model verification is performed on the risk prediction model, so that when it is determined that the first model verification fails, the model parameters of the risk prediction model are updated.

[0445] Optionally, the transaction risk prediction device 2300 further includes a second testing unit (not shown), which is specifically configured to:

[0446] If it is determined that the first model verification passes, obtaining second transaction features and second transaction tags of a plurality of second test transactions, wherein a transaction time difference between the second test transaction and the first test transaction satisfies a predetermined condition;

[0447] For each second test transaction, inputting the second transaction feature of the second test transaction into the risk prediction model to obtain a second transaction risk score of the second test transaction;

[0448] Based on the second transaction risk scores and second transaction labels of multiple second test transactions, a second model verification is performed on the risk prediction model, so that when it is determined that the second model verification fails, the model parameters of the risk prediction model are updated, or when the second model verification passes, the risk prediction model is deployed online.

[0449] Optionally, performing a second model verification on the risk prediction model based on the second transaction risk scores and the second transaction labels of the plurality of second test transactions includes:

[0450] determining a confusion matrix of the risk prediction model based on the second transaction risk scores, the second transaction labels, and a predetermined threshold for the plurality of second test transactions;

[0451] According to the first validation rule, the risk prediction model is subjected to the first sub-validation based on the confusion matrix;

[0452] According to the second validation rule, a second sub-validation is performed on the risk prediction model based on the confusion matrix.

[0453] Optionally, according to the first verification rule, based on the confusion matrix, a first sub-verification is performed on the risk prediction model, including:

[0454] Based on the confusion matrix, determine the information ratio, maximum cumulative transaction risk score difference, and feature discrimination accuracy of the risk prediction model;

[0455] The risk prediction model was first sub-validated based on the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy.

[0456] Optionally, according to the second verification rule, based on the confusion matrix, a second sub-verification is performed on the risk prediction model, including:

[0457] Determining, based on the confusion matrix, a first number of second test transactions whose second transaction risk scores are greater than or equal to a predetermined threshold, a second number of second test transactions whose second transaction label is 1, and a third number of second test transactions whose second transaction risk scores are greater than or equal to the predetermined threshold and whose second transaction label is 1;

[0458] According to the second verification rule, determining the transaction coverage of the risk prediction model based on the quotient of the third number and the second number, and determining the strategy cost performance of the risk prediction model based on the quotient of the first number and the third number;

[0459] The risk prediction model is secondly verified based on transaction coverage and strategy cost-effectiveness.

[0460] Optionally, the predetermined threshold is determined by:

[0461] Determine the current application scenario type of the risk prediction model;

[0462] Based on the current application scenario type, a predetermined threshold is determined from a plurality of candidate thresholds.

[0463] Optionally, the predetermined regularization coefficient is determined by:

[0464] Determine the sparsity of the sample reconstruction coefficient vector;

[0465] Based on the sparsity, a predetermined regularization coefficient is determined from among a plurality of candidate regularization coefficients.

[0466] Optionally, there are multiple risk prediction models, and each candidate transaction type corresponds to a risk prediction model;

[0467] The prediction unit 2320 is used to:

[0468] Determining a target transaction type of the target transaction from multiple candidate transaction types based on target transaction characteristics;

[0469] The target transaction features are input into the risk prediction model corresponding to the target transaction type to obtain the target transaction risk score of the target transaction.

[0470] Optionally, the predetermined condition is determined by:

[0471] Determining the training accuracy of the risk prediction model and the total number of predetermined transactions for the second test transaction;

[0472] Based on the training accuracy and the total number of predetermined transactions, an extreme value of the transaction time difference between the second test transaction and the first test transaction is determined, and a predetermined condition is determined based on the extreme value of the time difference.

[0473] Optionally, the transaction risk prediction device 2300 further includes a collection unit (not shown), which is specifically configured to:

[0474] Determine the data collection conditions for the risk prediction model;

[0475] If it is determined that the model training and model operation of the risk prediction model meet the data collection conditions, the log data of the risk prediction model is obtained to update the model based on the log data.

[0476] Reference Figure 24 , Figure 24 The following is a block diagram of the structure of a terminal for implementing the transaction risk prediction method according to an embodiment of the present disclosure. The terminal includes: a radio frequency (RF) circuit 2410, a memory 2415, an input unit 2430, a display unit 2440, a sensor 2450, an audio circuit 2460, a wireless fidelity (WiFi) module 2470, a processor 2480, and a power supply 2490. It will be understood by those skilled in the art that Figure 24 The terminal structure shown does not constitute a limitation on the mobile phone or computer, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0477] The RF circuit 2410 may be used for receiving and sending signals during information transmission or calls. In particular, after receiving downlink information from the base station, it is sent to the processor 2480 for processing. In addition, the designed uplink data is sent to the base station.

[0478] The memory 2415 may be used to store software programs and modules. The processor 2480 executes various functional applications and data processing of the target terminal by running the software programs and modules stored in the memory 2415 .

[0479] The input unit 2430 may be configured to receive input digital or character information and generate key signal input related to the setting and function control of the target terminal. Specifically, the input unit 2430 may include a touch panel 2431 and other input devices 2432.

[0480] The display unit 2440 may be configured to display input information or provided information and various menus of the target terminal. The display unit 2440 may include a display panel 2441.

[0481] The audio circuit 2460 , the speaker 2461 , and the microphone 2462 may provide an audio interface.

[0482] In this embodiment, the processor 2480 included in the terminal can execute the transaction risk prediction method of the previous embodiment.

[0483] The terminals of the embodiments of the present disclosure include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The embodiments of the present invention can be applied to various scenarios, including but not limited to data security, blockchain, data storage, information technology, etc.

[0484] Figure 25 A block diagram of the structure of a portion of a server for implementing the transaction risk prediction method of an embodiment of the present disclosure. The server may have relatively large differences due to different configurations or performance, and may include one or more central processing units (CPUs) 2522 (for example, one or more processors) and memories 2532, and one or more storage media 2530 (for example, one or more mass storage devices) for storing application programs 2542 or data 2544. Among them, the memories 2532 and the storage media 2530 can be temporary storage or permanent storage. The program stored in the storage medium 2530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 2522 can be configured to communicate with the storage medium 2530 to execute a series of instruction operations in the storage medium 2530 on the server.

[0485] The server may also include one or more power supplies 2526, one or more wired or wireless network interfaces 2550, one or more input and output interfaces 2558, and / or one or more operating systems 2541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0486] The central processor 2522 in the server can be used to execute the transaction risk prediction method of the embodiment of the present disclosure.

[0487] The embodiments of the present disclosure also provide a computer-readable storage medium, which is used to store program code, and the program code is used to execute the transaction risk prediction method of each of the aforementioned embodiments.

[0488] The present disclosure also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the transaction risk prediction method.

[0489] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "comprises" and "comprising," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0490] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0491] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0492] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0493] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0494] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0495] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0496] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0497] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A transaction risk prediction method, characterized in that: The method comprises: Obtain target transaction characteristics of the target transaction; Inputting the target transaction characteristics of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction, and performing transaction processing on the target transaction based on the target transaction risk score; Among them, the risk prediction model is obtained by training a preset logistic regression model based on the comparison of sample transaction risk scores of multiple sample transactions and sample transaction labels of the sample transactions; the sample transaction risk score is obtained by inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model; the sample reconstruction coefficient vector is predicted based on the sample transaction association data of the sample transaction, and the pre-set reference risk transaction characteristics and reference security transaction characteristics.

2. The transaction risk prediction method according to claim 1, characterized in that: The risk prediction model includes a reconstruction function and a prediction network; Inputting the target transaction characteristics of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction includes: Calculating a target reconstruction coefficient vector of the target transaction based on the reference risk transaction feature, the reference safety transaction feature, and the target transaction feature by the reconstruction function; The target reconstruction coefficient vector is input into the prediction network for probability calculation to obtain the target transaction risk score.

3. The transaction risk prediction method according to claim 1, characterized in that: The performing transaction processing on the target transaction based on the target transaction risk score includes: Invoking a plurality of predetermined candidate strategies and triggering conditions corresponding to each of the predetermined candidate strategies; If it is determined that the target transaction risk score and the target transaction characteristics meet the triggering condition of a target policy among the plurality of predetermined candidate policies, transaction processing is performed on the target transaction based on the target policy.

4. The transaction risk prediction method according to claim 1, characterized in that: The risk prediction model is trained in the following way: Obtaining sample transaction association data and sample transaction tags of a plurality of the sample transactions; For each of the sample transactions, based on a mapping result of the sample transaction-related data, obtaining a sample transaction feature of the sample transaction; For each of the sample transactions, determining a sample reconstruction coefficient vector of the sample transaction based on the sample transaction feature, the reference risk transaction feature, and the reference safety transaction feature; For each of the sample transactions, input the sample reconstruction coefficient vector into the logistic regression model, and output a sample transaction risk score of the sample transaction; Based on the comparison of the sample transaction risk scores of the plurality of sample transactions and the sample transaction labels, the model parameters of the logistic regression model are adjusted to obtain the risk prediction model.

5. The transaction risk prediction method according to claim 4, characterized in that: The sample transaction associated data includes the sample transaction content of the sample transaction and the sample transaction associated object; The acquiring, for each of the sample transactions, sample transaction features of the sample transaction based on a mapping result of the sample transaction associated data, includes: Obtaining a sample transaction content feature of the sample transaction based on a mapping result of the sample transaction content; Obtaining a sample transaction-associated object feature of the sample transaction based on a mapping result of the sample transaction-associated object; The sample transaction content feature and the sample transaction associated object feature are integrated into the sample transaction feature.

6. The transaction risk prediction method according to claim 4, characterized in that: The step of determining, for each of the sample transactions, a sample reconstruction coefficient vector of the sample transaction based on the sample transaction feature, the reference risk transaction feature, and the reference safety transaction feature, includes: generating a predetermined transaction feature matrix based on the reference risk transaction feature and the reference safety transaction feature; Determining an objective function based on the predetermined transaction feature matrix and a predetermined regularization coefficient; For each of the sample transactions, the sample transaction characteristics are substituted into the objective function, and the regularization term parameter that minimizes the output result of the objective function is determined as the sample reconstruction coefficient vector of the sample transaction.

7. The transaction risk prediction method according to claim 4, characterized in that: The logistic regression model includes an activation function, a first parameter, and a second parameter; Inputting the sample reconstruction coefficient vector into the logistic regression model and outputting the sample transaction risk score of the sample transaction includes: Inputting the sample reconstruction coefficient vector into the logistic regression model, and obtaining a first calculation result based on a transposed result of the first parameter and a product of the sample reconstruction coefficient vector; Probability scoring is performed on the sum of the first calculation result and the second parameter based on the activation function to obtain a sample transaction risk score of the sample transaction.

8. The transaction risk prediction method according to claim 4, characterized in that: The step of adjusting the model parameters of the logistic regression model based on the comparison of the sample transaction risk scores of the plurality of sample transactions with the sample transaction labels to obtain the risk prediction model includes: For each of the sample transactions, calculating a sub-loss function based on the sample transaction risk score and the sample transaction label; Obtaining a total loss function based on the sum of the multiple sub-loss functions; The model parameters of the logistic regression model are adjusted based on the total loss function to obtain the risk prediction model.

9. The transaction risk prediction method according to claim 8, characterized in that: The sample transactions include positive sample transactions and negative sample transactions; The calculating of a sub-loss function for each of the sample transactions based on the sample transaction risk score and the sample transaction label includes: If it is determined that the sample transaction is a positive sample transaction, determining the sub-loss function based on the negative logarithm of the risk score of the sample transaction; If it is determined that the sample transaction is a negative sample transaction, the sub-loss function is determined based on the negative logarithm of the difference between 1 and the risk score of the sample transaction.

10. The transaction risk prediction method according to claim 4, characterized in that: After adjusting the model parameters of the logistic regression model to obtain the risk prediction model, the method further includes: Obtaining first transaction features and first transaction tags of a plurality of first test transactions; For each of the first test transactions, inputting the first transaction feature of the first test transaction into the risk prediction model to obtain a first transaction risk score of the first test transaction; Based on the first transaction risk scores of the plurality of first test transactions and the first transaction labels, a first model verification is performed on the risk prediction model, so as to update the model parameters of the risk prediction model when it is determined that the first model verification fails.

11. The transaction risk prediction method according to claim 10, characterized in that: After performing the first model verification on the risk prediction model, the method further includes: If it is determined that the first model verification passes, obtaining second transaction features and second transaction tags of a plurality of second test transactions, wherein a transaction time difference between the second test transaction and the first test transaction satisfies a predetermined condition; For each second test transaction, inputting the second transaction feature of the second test transaction into the risk prediction model to obtain a second transaction risk score of the second test transaction; Based on the second transaction risk scores of multiple second test transactions and the second transaction labels, the risk prediction model is subjected to a second model verification, so as to update the model parameters of the risk prediction model when it is determined that the second model verification fails, or to deploy the risk prediction model online when the second model verification passes.

12. The transaction risk prediction method according to claim 11, characterized in that: The performing a second model verification on the risk prediction model based on the second transaction risk scores of the plurality of second test transactions and the second transaction labels includes: determining a confusion matrix of the risk prediction model based on the second transaction risk scores of the plurality of second test transactions, the second transaction labels, and a predetermined threshold; According to a first validation rule, based on the confusion matrix, performing a first sub-validation on the risk prediction model; According to the second verification rule, based on the confusion matrix, a second sub-verification is performed on the risk prediction model.

13. The transaction risk prediction method according to claim 12, characterized in that: The step of performing a first sub-verification on the risk prediction model based on the confusion matrix according to the first verification rule includes: Determining, based on the confusion matrix, an information ratio, a maximum cumulative transaction risk score difference, and a feature discrimination accuracy of the risk prediction model; Based on the information ratio, the maximum cumulative transaction risk score difference, and the feature discrimination accuracy, a first sub-verification is performed on the risk prediction model.

14. The transaction risk prediction method according to claim 12, characterized in that: The step of performing a second sub-verification on the risk prediction model based on the confusion matrix according to the second verification rule includes: determining, based on the confusion matrix, a first number of second test transactions for which the second transaction risk score is greater than or equal to a predetermined threshold, a second number of second test transactions for which the second transaction label is 1, and a third number of second test transactions for which the second transaction risk score is greater than or equal to the predetermined threshold and the second transaction label is 1; Determining, according to the second verification rule, a transaction coverage of the risk prediction model based on a quotient of the third number and the second number, and determining a strategy cost-effectiveness of the risk prediction model based on a quotient of the first number and the third number; A second sub-verification is performed on the risk prediction model based on the transaction coverage and the strategy cost-effectiveness.

15. The transaction risk prediction method according to claim 12, characterized in that: The predetermined threshold is determined by: Determining the current application scenario type of the risk prediction model; The predetermined threshold is determined from a plurality of candidate thresholds based on the current application scenario type.

16. The transaction risk prediction method according to claim 6, characterized in that: The predetermined regularization coefficient is determined by: Determining the sparsity of the sample reconstruction coefficient vector; The predetermined regularization coefficient is determined from a plurality of candidate regularization coefficients based on the sparsity.

17. A transaction risk prediction device, characterized in that: The device comprises: an acquisition unit, configured to acquire target transaction characteristics of a target transaction; a prediction unit, configured to input a target transaction feature of the target transaction into a risk prediction model to obtain a target transaction risk score of the target transaction, and perform transaction processing on the target transaction based on the target transaction risk score; Among them, the risk prediction model is obtained by training a preset logistic regression model based on the comparison of sample transaction risk scores of multiple sample transactions and sample transaction labels of the sample transactions; the sample transaction risk score is obtained by inputting the sample reconstruction coefficient vector of the sample transaction into the logistic regression model; the sample reconstruction coefficient vector is predicted based on the sample transaction association data of the sample transaction, and the pre-set reference risk transaction characteristics and reference security transaction characteristics.

18. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the transaction risk prediction method according to any one of claims 1 to 16 is implemented.

19. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the transaction risk prediction method according to any one of claims 1 to 16 is implemented.

20. A computer program product, comprising a computer program, wherein the computer program is read and executed by a processor of a computer device, so that the computer device executes the transaction risk prediction method according to any one of claims 1 to 16.