Method and apparatus for training risk identification model

The labeled sample set is enhanced and soft label generation is generated through interpolation, which solves the problem of the utilization of labeled samples in the risk identification model and improves the recognition accuracy and performance of the model.

WO2025139251A1PCT designated stage expired Publication Date: 2025-07-03TSINGHUA UNIVERSITY +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/125980
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-10-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, there are a large number of suspicious transactions without labels in real application scenarios, resulting in a degradation in the performance of supervised learning models when identifying risk transactions, and the imbalance of sample categories in the labeled sample set and the generalization problem of unlabeled sample distribution is difficult to effectively utilize.

Method used

The sample enhancement of the annotated sample set is performed through interpolation, soft labels are generated, and risk identification models are trained in combination with the annotated and unlabeled sample sets. The soft labels and hard labels generated by the interpolation method are used for model training to alleviate the problem of sample category imbalance and distribution generalization.

Benefits of technology

Effective utilization of unlabeled suspicious transaction data has improved the performance of the risk identification model, alleviated the problems of sample category imbalance and distribution generalization, and improved the accuracy of the model identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125980_03072025_PF_FP_ABST
    Figure CN2024125980_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present specification relate to a method and apparatus for training a risk identification model. The method comprises: first, acquiring a first sample set having hard labels and a second sample set having no label, wherein any sample set comprises transaction samples, and each hard label indicates whether a transaction is a risk transaction; then, performing sample enhancement on the first sample set on the basis of an interpolation method, and using the enhanced first sample set to perform training, so as to obtain a first model; subsequently, inputting, into the first model, transaction samples in a complete sample set formed by the first sample set and the second sample set, so as to obtain soft labels concerning risk prediction; finally, inputting the transaction samples in the first sample set into a second model, and determining a first loss on the basis of the hard labels; inputting the transaction samples in the complete sample set into the second model, and determining a second loss on the basis of the soft labels; and training the second model on the basis of a total predicted loss determined by means of the first loss and the second loss, wherein the second model is used for predicting whether each transaction is a risky transaction.
Need to check novelty before this filing date? Find Prior Art

Description

A method and device for training risk identification model

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 28, 2023, with application number 2023118452089 and application name “A method and device for training a risk identification model”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] One or more embodiments of this specification relate to the field of artificial intelligence, and in particular, to a method and apparatus for training a risk identification model. Background Art

[0003] With the booming development of mobile payments and e-commerce, transaction volumes on e-service platforms are increasing. However, at the same time, many risky users are exploiting e-payment methods to engage in fraudulent transactions, account theft, and other risky transactions, seriously infringing on the rights and interests of other ordinary users. Identifying and addressing potentially risky transactions is a crucial step in ensuring the security of service platforms and protecting user assets and transactions.

[0004] Currently, supervised machine learning technology is widely used in risky transaction detection. Supervised learning relies on labeled positive and negative examples to train the model. However, in real-world applications, there are often a large number of unlabeled suspicious transactions, which cannot be directly applied to supervised learning. Therefore, a method for training risk identification models is needed to effectively utilize these suspicious transaction samples and improve the overall performance of risk identification models.

[0005] Summary of the Invention

[0006] One or more embodiments of this specification describe a method and apparatus for training a risk identification model, aiming to improve the performance of the risk identification model by effectively utilizing unlabeled suspicious transaction data.

[0007] In a first aspect, a method for training a risk identification model is provided, comprising:

[0008] Obtaining a first sample set with hard labels and a second sample set without labels, either sample set including transaction samples, wherein the hard labels indicate whether the corresponding transactions are risky transactions;

[0009] Performing sample enhancement on the first sample set based on an interpolation method, and using the enhanced first sample set to train a first model;

[0010] Inputting each transaction sample in the total sample set consisting of the first sample set and the second sample set into the first model to obtain a soft label for each transaction sample regarding risk prediction;

[0011] Inputting transaction samples in the first sample set into a second model, and determining a first loss based on hard labels of each transaction sample;

[0012] Inputting transaction samples in the total sample set into a second model, and determining a second loss based on the soft label of each transaction sample;

[0013] The second model is trained based on a total predicted loss determined by the first loss and the second loss, and the second model is used to predict whether a transaction is a risky transaction.

[0014] In a possible implementation, the transaction samples included in the second sample set are transaction samples identified as suspicious transactions by the risk detection system during the transaction occurrence stage.

[0015] In a possible implementation, performing sample enhancement on the first sample set based on an interpolation method includes:

[0016] performing weighted summation on the sample features of the first transaction sample and the second transaction sample in the first sample set to obtain an interpolated sample feature;

[0017] Performing a weighted summation on the first hard label of the first transaction sample and the second hard label of the second transaction sample to obtain an interpolated label;

[0018] The interpolated sample features and the interpolated labels constitute enhanced samples, which are added to the first sample set.

[0019] In one possible implementation, using the enhanced first sample set to train a first model includes:

[0020] Determining a training loss based on a cross entropy loss between a predicted value of each sample in the enhanced first sample set and a corresponding hard label by the first model;

[0021] The first model is trained based on the training loss.

[0022] In one possible implementation, determining the first loss based on the hard label of each transaction sample includes:

[0023] The first loss is determined based on the cross entropy loss between the predicted value of each transaction sample and the corresponding hard label of the second model.

[0024] In one possible implementation, determining the second loss based on the soft label of each transaction sample includes:

[0025] A second loss is determined based on a cross entropy loss between a predicted value of each transaction sample and a corresponding soft label by the second model.

[0026] In one possible implementation, the first model and the second model include a softmax layer with a temperature coefficient.

[0027] In one possible implementation, the total prediction loss is determined by the following process:

[0028] The total predicted loss is determined based on a weighted sum of the first loss and the second loss, wherein a weight coefficient of the weighted sum includes the temperature coefficient.

[0029] In one possible implementation, each transaction sample in the total sample set consisting of the first sample set and the second sample set is input into the first model to obtain a soft label for each transaction sample regarding risk prediction, including:

[0030] Input each transaction sample into the first model to obtain an output result vector of the first model before the softmax layer with a temperature coefficient;

[0031] The output result vector is smoothed by the softmax layer with a temperature coefficient to obtain the soft label.

[0032] In a possible implementation, the hard label and the soft label are in the form of two-dimensional vectors; the values ​​of the two dimensions of the two-dimensional vectors respectively indicate the probability values ​​of the corresponding transaction being a normal transaction and a risky transaction.

[0033] In a second aspect, a device for training a risk identification model is provided, comprising:

[0034] an acquisition unit configured to acquire a first sample set with a hard label and a second sample set without a label, wherein either sample set includes a transaction sample, and the hard label indicates whether the corresponding transaction is a risky transaction;

[0035] A first model training unit is configured to perform sample enhancement on the first sample set based on an interpolation method, and train a first model using the enhanced first sample set;

[0036] a soft label generating unit configured to input each transaction sample in the total sample set consisting of the first sample set and the second sample set into the first model to obtain a soft label for each transaction sample regarding risk prediction;

[0037] a first loss determining unit configured to input transaction samples in the first sample set into a second model and determine a first loss based on a hard label of each transaction sample;

[0038] a second loss determining unit configured to input the transaction samples in the total sample set into a second model, and determine a second loss based on the soft label of each transaction sample;

[0039] The second model training unit is configured to train the second model based on the total predicted loss determined by the first loss and the second loss, where the second model is used to predict whether a transaction is a risky transaction.

[0040] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0041] In a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0042] The embodiments of this specification propose a method and device for training a risk identification model. Through sample enhancement based on interpolation and automatic labeling of suspicious transaction samples, they effectively utilize suspicious transactions, alleviate the problem of imbalance between positive and negative sample categories, and improve the performance of the risk identification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0044] FIG1 shows a schematic diagram of sample distribution according to an example;

[0045] FIG2 is a schematic diagram showing an implementation scenario of a method for training a risk identification model according to an embodiment;

[0046] FIG3 shows a flow chart of a method for training a risk identification model according to one embodiment;

[0047] FIG4 shows a schematic block diagram of an apparatus for training a risk identification model according to one embodiment. DETAILED DESCRIPTION

[0048] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0049] As mentioned earlier, supervised machine learning is currently widely used in risky transaction detection. Supervised learning relies on labeled positive and negative examples to train the model. However, in real-world applications, there are often a large number of unlabeled suspicious transactions, which cannot be directly applied to supervised learning.

[0050] In practice, after a transaction is initiated, it often first undergoes an online risk detection system. If it passes the test, it is considered a normal transaction and cleared by the system. If it fails the test, it is intercepted by the system and considered suspicious. Furthermore, for transactions identified as normal by the system, some normal transactions may be identified as risky transactions through user reports, verification by the service provider, and self-verification by the service provider. Normal transactions and risky transactions are labeled samples, representing positive and negative samples, respectively. Suspicious transactions intercepted by the system are unlabeled samples and may be either normal or risky transactions, requiring further verification.

[0051] After analyzing both labeled and unlabeled samples, the inventors discovered that since most users are ordinary and generally conduct normal transactions, only a small fraction of them engage in risky transactions. Therefore, the volume of risky transactions in the labeled samples is far smaller than that of normal transactions. This leads to an imbalance in the sample categories within the labeled sample set, making it difficult to directly train supervised machine learning models. Furthermore, some normal transactions contain risky transactions that have not been reported by users or detected by service providers, representing noise in the normal transactions.

[0052] For unlabeled samples, in addition to the problem that the samples themselves are unlabeled, the overall data distribution of unlabeled samples will also deviate from that of labeled samples, forming a characteristic clustered distribution. For example, Figure 1 shows a schematic diagram of sample distribution based on an example. As shown in Figure 1, normal transactions are circles, suspicious transactions are triangles, and risky transactions are forks. The distribution of suspicious transactions will deviate from normal transactions and risky transactions, which is called out-of-distribution (OOD). For a supervised model that is only trained on labeled samples, its performance will drop significantly when it recognizes data with a different distribution from the training data. It should be noted that Figure 1 only shows the case where the feature dimension of the sample is 2-dimensional. In actual applications, the feature dimension of the sample can be greater than 2-dimensional.

[0053] To summarize the above, if we refer to normal transaction samples as white samples and risky transaction samples as black samples, then in practical applications, the problems with labeled samples are class imbalance between black and white samples and white sample noise. Unlabeled samples can be referred to as gray samples. The problems with gray samples are unlabeled data and out-of-distribution generalization. In other words, gray samples can be defined as out-of-distribution unlabeled data. Due to the problems inherent in both black, white, and gray samples, it is difficult to use traditional supervised models to directly learn from them.

[0054] In order to solve the above problems, Figure 2 shows a schematic diagram of an implementation scenario of a method for training a risk identification model according to one embodiment. As shown in Figure 2, after a transaction is initiated, it undergoes preliminary identification by the risk detection system, and some transactions are released while others are intercepted by the system. Among the released transactions, some transactions are labeled as risky transactions (black samples) after user reports and verification by the service provider, or self-verification by the service provider, etc., while the remaining transactions are considered normal transactions (white samples). These labeled transaction samples form a black and white sample set and have corresponding hard labels. A hard label is a label with a binary value range, which can be a constant, for example, the hard label of a white sample is 1 and the hard label of a black sample is 0. Alternatively, a hard label can also be a two-dimensional vector, for example, the hard label of a white sample is (1,0) and the hard label of a black sample is (0,1). Transactions intercepted by the risk detection system form a gray sample set, and the transaction samples in this set have no labels.

[0055] Next, a risk identification model is trained based on the two sample sets. First, the black and white sample sets are augmented using interpolation. Model 1 is then trained based on the augmented black and white sample sets. This trained model 1 outputs the probability of a transaction being normal or risky. The probability can be a constant, representing the probability of a transaction being normal, or a two-dimensional vector, where the first dimension represents the probability of a transaction being normal and the second dimension represents the probability of a transaction being risky. Then, each transaction sample from the black and white sample sets and the gray sample set are input into the trained model 1 to obtain the probability corresponding to each transaction sample. This probability is called a soft label for the corresponding transaction sample, and the soft label is added to the corresponding sample set. Compared to hard labels, which are binary values ​​of 0-1, soft labels, as continuous variables, can contain more information. Finally, based on the hard and / or soft labels corresponding to each transaction sample in the two sample sets, a new model 2 is trained. Model 2 serves as the risk identification model and can better identify potentially risky transactions.

[0056] The following describes the specific implementation steps of the above-mentioned method for training a risk identification model in conjunction with specific embodiments. Figure 3 shows a flowchart of the method for training a risk identification model according to one embodiment. The execution subject of the method can be any platform, server, or device cluster with computing and processing capabilities. As shown in Figure 3, the method includes at least the following steps: Step 302: obtaining a first sample set with hard labels and a second sample set without labels, either sample set including transaction samples, wherein the hard labels indicate whether the corresponding transactions are risky transactions; Step 304: performing sample enhancement on the first sample set using an interpolation method, and training a first model using the enhanced first sample set; Step 306: inputting each transaction sample from the total sample set consisting of the first and second sample sets into the first model to obtain a soft label for each transaction sample regarding risk prediction; Step 308: inputting the transaction samples from the first sample set into the second model to determine a first loss based on the hard labels of each transaction sample; Step 310: inputting the transaction samples from the total sample set into the second model to determine a second loss based on the soft labels of each transaction sample; and Step 312: training a second model based on the total predicted loss determined by the first and second losses, wherein the second model is used to predict whether a transaction is risky. The detailed execution process of each of these steps is described below.

[0057] First, in step 302 , a first sample set with hard labels and a second sample set without labels are obtained, where either sample set includes transaction samples, and the hard labels indicate whether the corresponding transactions are risky transactions.

[0058] The first sample set can be the aforementioned labeled black and white sample set, which includes normal transactions (white samples) and risky transactions (black samples). The second sample set can be the aforementioned gray sample set, which includes suspicious transactions (gray samples). Formally, the first sample set can be denoted as D = {(x1, y1), (x2, y2), …, (x n ,y n )}, y represents the sample feature of the i-th transaction sample, and d is the feature dimension of the transaction sample. i is x i The labeling information is a binary hard label with a value range of {0,1}, indicating x i Is it a risky transaction?

[0059] In one embodiment, the hard label can be in the form of a two-dimensional vector, where the values ​​of the two dimensions indicate the probability of the corresponding transaction being a normal transaction or a risky transaction, respectively. The hard label of a white sample is (0, 1), and the hard label of a black sample is (1, 0).

[0060] Then, the second sample set can be recorded as D′={x′1, x′2,…, x′ m}, In a possible implementation, m<<n, that is, the number of samples in the second sample set is much smaller than that in the first sample set.

[0061] In one embodiment, the transaction samples included in the second sample set D′ are those identified as suspicious transactions by the risk detection system during the transaction phase. The risk detection system can be rule-based or machine learning-based, and this is not limited here. Since the suspicious transaction samples (gray samples) in the second sample set D′ have already been identified by the risk detection system, their data distribution is more biased towards black samples. When they are subsequently soft-labeled, the soft label will be more likely to indicate that they are black samples.

[0062] Then, in step 304, sample enhancement is performed on the first sample set based on an interpolation method, and the enhanced first sample set is used to train a first model.

[0063] Specifically, the first sample set is enhanced based on the interpolation method: the sample features of the first transaction sample and the second transaction sample in the first sample set are weightedly summed to obtain interpolated sample features; the first hard label of the first transaction sample and the second hard label of the second transaction sample are weightedly summed to obtain an interpolated label; the interpolated sample features and the interpolated labels constitute an enhanced sample and are added to the first sample set.

[0064] The sample feature of the first transaction sample can be recorded as x i , the corresponding first hard label can be recorded as y i , the sample feature of the second transaction sample can be recorded as x i , and its corresponding second hard label can be recorded as y i . Then calculate the interpolation sample features It can be shown as formula (1):

[0065] Compute interpolated labels It can be shown as formula (2):

[0066] Among them, λ is a hyperparameter with a value range between (0,1).

[0067] Then, the interpolated sample features and interpolation labels Composition enhancement sample Added to the first sample set D.

[0068] xi and x j They can be all white samples, all black samples, or one white sample and one black sample. i and x j When there is a white sample and a black sample, the interpolation sample features obtained based on interpolation The data distribution will be far away from the sample distribution of white samples and black samples, and to some extent close to the distribution of gray samples.

[0069] Repeat the above sample enhancement steps multiple times, and add the enhanced samples obtained in each time to the first sample set D to obtain the enhanced first sample set D * .

[0070] Then, use the enhanced first sample set D * The first model is trained to obtain a first model. The first model can be Model 1 described above, or any binary classification model, such as a multilayer perceptron, without limitation. In one embodiment, the first model includes a softmax layer with a temperature coefficient t. The temperature coefficient t is a hyperparameter greater than 1 and is used to smooth the model output.

[0071] Specifically, based on the first model, for the enhanced first sample set D * The cross entropy loss between the predicted value of each sample and the corresponding hard label determines the training loss L t ; Then, based on the training loss L t , train the first model.

[0072] Determine the training loss L t It can be shown as formula (3): l(f1(x i ),y i )=y i1 log(f1(x i )1)+y i2 log(f1(x i )twenty four)

[0073] Among them, f1(x i ) represents x i The output vector of the first model is used as input. f1(x i )1 represents the value of the first dimension of the output vector, f1(x i )2 represents the value of the second dimension of the output vector. i1 is x i Hard tag y i The value of the first dimension, y i2 is x i Hard tag y i The value of the second dimension.

[0074] Based on the training loss L t After training the first model using the gradient descent method, a trained first model is obtained.

[0075] Next, in step 306, each transaction sample in the total sample set consisting of the first sample set and the second sample set is input into the first model to obtain a soft label for each transaction sample regarding risk prediction.

[0076] The soft label can be in the form of a two-dimensional vector, where the values ​​of the two dimensions indicate the probability of the corresponding transaction being a normal transaction or a risky transaction. i Input it into the first trained model to get the output result vector (logits) z of the first model before the softmax layer with temperature coefficient t i =(z i1 ,z i2 ), and then the soft label is obtained by smoothing the softmax layer with temperature coefficient t As shown in formula (5):

[0077] Among them, φ() corresponds to the calculation process of the softmax layer with a temperature coefficient t.

[0078] Each transaction sample x i Soft label Add to the first sample set D to obtain the expanded first sample set The expanded first sample set D + Each transaction sample x in i Has the corresponding hard label y i and soft labels

[0079] Each transaction sample x′ in the second sample set D′ i Input it into the first trained model to obtain the output result vector z′ of the first model before the softmax layer with temperature coefficient t i =(z′ i1 ,z′ i2 ), and then the soft label is obtained by smoothing the softmax layer with temperature coefficient t As shown in formula (7):

[0080] Each transaction sample x′ i Soft label Add to the second sample set D′ to obtain the expanded second sample set The expanded second sample set D′ + Each transaction sample x′ i Have corresponding soft labels

[0081] Then, in step 308 , the transaction samples in the first sample set are input into the second model, and a first loss is determined based on the hard labels of the respective transaction samples.

[0082] The second model can be the aforementioned model 2, or any binary classification model, such as a multi-layer perceptron, which is not limited here. In one embodiment, the second model includes a softmax layer with a temperature coefficient t.

[0083] Specifically, the first loss L is determined based on the cross entropy loss between the predicted value of each transaction sample and the corresponding hard label of the second model. h .

[0084] Determine the first loss L h It can be shown as formula (9): l(f2(x i ),y i )=y i1 log(f2(x i )1)+y i2 log(f2(x i )2) (10)

[0085] Among them, f2(x i ) represents x i The output vector of the second model when used as input. f2(x i )1 represents the value of the first dimension of the output vector, f2(x i )2 represents the value of the second dimension of the output vector. i Input into the second model, get the output result (logits) of the second model before the softmax layer with temperature coefficient t, and then get f2(x i ).

[0086] Furthermore, in step 310 , the transaction samples in the total sample set are input into the second model, and the second loss is determined based on the soft label of each transaction sample.

[0087] Specifically, based on the cross entropy loss between the predicted value of each transaction sample and the corresponding soft label of the second model, the second loss L is determined. s .

[0088] Determine the second loss L s It can be shown as formula (11):

[0089] Among them, f2(x′ i ) represents x′ i The output vector of the second model when used as input. i )1 represents the value of the first dimension of the output vector, f2(x′ i )2 represents the value of the second dimension of the output vector. i Input into the second model, get the output result vector (logits) of the second model before the softmax layer with temperature coefficient t, and then get f2(x′) through the softmax layer with temperature coefficient t i ). is x i Soft label The value of the first dimension of is x i Soft label The value of the second dimension. is x′ i Soft label The value of the first dimension of is x′ i Soft label The value of the second dimension.

[0090] Finally, in step 312, based on the first loss L h and the second loss L s The determined total prediction loss L is used to train the second model, where the second model is used to predict whether a transaction is a risky transaction.

[0091] In one embodiment, the total predicted loss is determined based on a weighted sum of the first loss and the second loss, wherein the weight coefficient of the weighted sum includes the temperature coefficient. The total predicted loss L can be determined as shown in formula (14): L = (1-α) L h +αt 2 L s (14)

[0092] Where α is a parameter that controls the weights of the two losses. The second model is trained using gradient descent based on the total prediction loss L to obtain the trained second model. The second model is a risk identification model used to predict whether a transaction is risky.

[0093] In summary, the method for training a risk identification model described in the embodiment of this specification, based on the interpolation method and the specific model training method, effectively alleviates the problems mentioned at the beginning. Specifically, the embodiment of this specification performs data enhancement on the black and white sample sets through the interpolation method, which can deal with the distribution out-generalization problem existing in unlabeled samples. After using the first model to soft-label the gray samples, since the gray samples have undergone a round of recognition by the risk detection system, they will be more biased towards black samples in data distribution, which can effectively expand the number of black samples and deal with the imbalance of black and white sample categories in labeled samples. Through the specific training steps of the risk identification model (second model), the model can be better trained with labeled samples and unlabeled samples to improve the performance of the model, and deal with the white sample noise problem in labeled samples and the unlabeled problem in unlabeled samples. By introducing the temperature coefficient t, the second model can better learn the knowledge in the first model.

[0094] According to another embodiment, a device for training a risk identification model is also provided. FIG4 shows a schematic block diagram of a device for training a risk identification model according to one embodiment. The device can be deployed in any device, platform, or device cluster with computing and processing capabilities. As shown in FIG4 , the device 400 includes:

[0095] An acquisition unit 401 is configured to acquire a first sample set with a hard label and a second sample set without a label, wherein either sample set includes a transaction sample, and the hard label indicates whether the corresponding transaction is a risky transaction;

[0096] A first model training unit 402 is configured to perform sample enhancement on the first sample set based on an interpolation method, and train a first model using the enhanced first sample set;

[0097] The soft label generating unit 403 is configured to input each transaction sample in the total sample set consisting of the first sample set and the second sample set into the first model to obtain a soft label for each transaction sample regarding risk prediction;

[0098] A first loss determining unit 404 is configured to input the transaction samples in the first sample set into the second model and determine a first loss based on the hard labels of the respective transaction samples;

[0099] A second loss determining unit 405 is configured to input the transaction samples in the total sample set into a second model and determine a second loss based on the soft label of each transaction sample;

[0100] The second model training unit 406 is configured to train the second model based on the total predicted loss determined by the first loss and the second loss, where the second model is used to predict whether a transaction is a risky transaction.

[0101] According to another embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any one of the above embodiments.

[0102] According to yet another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the above embodiments is implemented.

[0103] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0104] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0105] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0106] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program instructing the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0107] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for training a risk identification model, comprising: Obtaining a first sample set with hard labels and a second sample set without labels, where any sample set includes transaction samples, and the hard labels indicate whether the corresponding transactions are risk transactions; Performing sample augmentation on the first sample set based on interpolation method, and training a first model using the augmented first sample set; Inputting each transaction sample in the total sample set composed of the first sample set and the second sample set into the first model to obtain soft labels for each transaction sample regarding risk prediction; Inputting the transaction samples in the first sample set into a second model, and determining a first loss based on the hard labels of each transaction sample; Inputting the transaction samples in the total sample set into the second model, and determining a second loss based on the soft labels of each transaction sample; Training the second model based on the total prediction loss determined by the first loss and the second loss, where the second model is used to predict whether a transaction is a risk transaction.

2. The method according to claim 1, wherein The transaction samples included in the second sample set are the transaction samples identified as suspicious transactions by the risk detection system during the transaction occurrence stage.

3. The method according to claim 1, performing sample augmentation on the first sample set based on interpolation method, comprising: Performing weighted summation on the sample features of a first transaction sample and a second transaction sample in the first sample set to obtain interpolated sample features; Performing weighted summation on the first hard label of the first transaction sample and the second hard label of the second transaction sample to obtain an interpolated label; Forming an augmented sample with the interpolated sample features and the interpolated label, and adding it to the first sample set.

4. The method according to claim 1, training a first model using the augmented first sample set, comprising: Determining a training loss based on the cross-entropy loss between the predicted value of each sample in the augmented first sample set by the first model and the corresponding hard label; Training the first model based on the training loss.

5. The method according to claim 1, determining a first loss based on the hard labels of each transaction sample, comprising: Determining a first loss based on the cross-entropy loss between the predicted value of each transaction sample by the second model and the corresponding hard label.

6. The method according to claim 1, determining a second loss based on the soft labels of each transaction sample, comprising: Determining a second loss based on the cross-entropy loss between the predicted value of each transaction sample by the second model and the corresponding soft label.

7. The method according to claim 1, wherein The first model and the second model include softmax layers with temperature coefficients.

8. The method according to claim 7, the total prediction loss is determined through the following process: Determining the total prediction loss based on the weighted sum of the first loss and the second loss, and the weight coefficients of the weighted sum include the temperature coefficient.

9. The method according to claim 7, inputting each transaction sample in the total sample set composed of the first sample set and the second sample set into the first model to obtain soft labels for each transaction sample regarding risk prediction, comprising: Input each transaction sample into the first model to obtain an output result vector of the first model before the softmax layer with a temperature coefficient; Smooth the output result vector through the softmax layer with a temperature coefficient to obtain the soft label.

10. The method according to claim 1, wherein The hard label and the soft label are in the form of two-dimensional vectors; the values of the two dimensions of the two-dimensional vector respectively indicate the probability values of the corresponding transaction being a normal transaction and a risky transaction.

11. An apparatus for training a risk identification model, comprising: An acquisition unit configured to acquire a first sample set with hard labels and a second sample set without labels, where any sample set includes transaction samples, and the hard label indicates whether the corresponding transaction is a risky transaction; A first model training unit configured to perform sample augmentation on the first sample set based on interpolation and use the augmented first sample set to train a first model; A soft label generation unit configured to input each transaction sample in the total sample set composed of the first sample set and the second sample set into the first model to obtain soft labels for each transaction sample regarding risk prediction; A first loss determination unit configured to input the transaction samples in the first sample set into a second model and determine a first loss based on the hard labels of each transaction sample; A second loss determination unit configured to input the transaction samples in the total sample set into the second model and determine a second loss based on the soft labels of each transaction sample; A second model training unit configured to train the second model based on the total prediction loss determined by the first loss and the second loss, where the second model is used to predict whether a transaction is a risky transaction.

12. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-10.

13. A computing device, comprising a memory and a processor, wherein, An executable code is stored in the memory. When the processor executes the executable code, the method according to any one of claims 1-10 is implemented.

Citation Information

Patent Citations

  • Risk transaction identification model training method, risk transaction identification method and device

    CN114912549A

  • Personalized human body action recognition method based on knowledge distillation

    CN116844225A

  • Depth neural network confidence coefficient calibration method and device, equipment and storage medium

    CN116956013A

  • Method and device for training risk identification model

    CN117743856A

  • Method and system with neural network model updating

    US20210034971A1