An e-commerce fraud identification method and system based on continuous learning
Patent Information
- Application Number
- CN202310972921.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-08-03
AI Technical Summary
由于保存数据需要占用大量存储空间,模型在训练时通常只使用最新一段时间内的数据,所以其并不能够充分接触和学习到历史风险点的特征
[0040]本发明设计的基于持续学习的电商欺诈识别方法,将持续学习框架引入到欺诈识别模型中去。对商家行为特征、商品介绍文本特征以及聊天文本数据特征的组合能够充分挖掘出交易过程所包含的信息。温度调节机制可以平滑模型训练时的输出分数,增加分布的熵,从而让模型获取更多信息。通过知识蒸馏的方法可以使新模型参数在更新过程中受到线上模型参数的引导,令其在推理过程中参照旧模型对样本的处理方式,这是保留历史知识的一种手段;而样本重演方法在占用有限额外存储资源的情况下,让模型直接接触到历史数据信息。以上两点可以提高新模型对历史风险点的识别准确率,有效缓解模型更新导致的灾难性遗忘问题。本发明中的持续学习框架可以通过简单调整应用于其它实际场景,使得本发明具备良好的通用性。
Smart Images

Figure CN117114705B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data mining technology, specifically relating to a method and system for identifying e-commerce fraud based on continuous learning. Background Technology
[0002] With the development of the internet, information accumulates continuously online, including a large amount of illegal information. In the e-commerce sector, fraudulent activities such as false product descriptions, selling counterfeit goods, and even failing to deliver goods are rampant, necessitating strict supervision of merchants and the transaction process. Currently, companies invest significant manpower in reviewing every piece of information. To save on labor costs and reduce review time, deep learning models, especially BERT-based natural language processing models, are used to uncover potential risks in merchant behavior data. However, existing methods have several problems. In the struggle between unscrupulous merchants and platform regulation, fraudulent methods are constantly evolving. The model currently in use (referred to as the "online model" or "old model") lacks the ability to identify the latest fraudulent methods, thus requiring updates using the latest fraud cases (referred to as the "new model"). Since storing data requires significant storage space, models typically only use data from the most recent period during training, thus failing to fully access and learn the characteristics of historical risk points. When historical fraudulent methods reappear in the future, currently running online models become unable to effectively identify them, leading to the neglect of a large amount of fraudulent information and a decrease in the overall classification accuracy of the model. This is the catastrophic forgetting problem that deep learning models face when data distribution changes. In real-world applications, a large amount of e-commerce-related data is generated daily (samples corresponding to the latest fraud methods are called "new risk point data," risk samples that have been learned by the model are called "old risk point data," and there are also a large number of risk-free samples). Retaining all historical data and using it for training would consume enormous computing resources and time, which is impractical. Therefore, there is an urgent need to design a technical solution that can overcome the above-mentioned shortcomings, so that the model can effectively identify new fraudulent methods after updates, while also retaining the learned historical knowledge to the greatest extent, thereby improving the overall risk identification capability of the model. Summary of the Invention
[0003] The purpose of this invention is to provide an e-commerce fraud identification method and device based on continuous learning, which can alleviate the catastrophic forgetting problem of the model as much as possible, enhance the model's ability to remember historical risk points after updates, and improve the overall identification performance of the model on both new and old risk points.
[0004] The e-commerce fraud detection method based on continuous learning provided by this invention includes the following specific steps:
[0005] (1) Sample feature extraction: The text information generated during the merchant's transaction process is encoded using the existing pre-trained vocabulary (a set of vector representations corresponding to words) and the text feature extractor (here, the encoder module in the Transformer model is selected), and then concatenated with the merchant's behavioral features to obtain sample features.
[0006] (2) Sample risk identification; The sample features learned by the model are scored by a binary classifier, and an appropriate threshold is set to obtain the final classification result;
[0007] (3) Model iteration based on continuous learning; by using knowledge distillation and sample replay methods, the parameters of the new model are brought closer to the parameters of the online model during the training process, so as to retain the ability of the new model to identify historical risks as much as possible while learning the feature information of new risk points.
[0008] Further, the sample feature extraction in step (1) specifically includes: extracting merchant behavior features, extracting features from product description text information, and extracting features from merchant product descriptions and consumer chat information; finally, concatenating various feature vectors and performing nonlinear transformation.
[0009] Specifically, for the merchant behavior data wide_x, two fully connected neural networks f1 and f2 and a ReLU activation function are used to extract the behavioral features wide_output:
[0010] wide_output=f2.Relu(f1(wide_x))) (1)
[0011] For text data such as product descriptions and chat messages between buyers and sellers, feature extraction is mainly performed in two ways: first, using the Word2vec model to convert the text into word vectors; second, directly obtaining the vector representation of each word from an existing pre-trained vocabulary using word indices. Based on this, positional encoding (position_embeddings) is added to the word vectors and encoded by an encoder consisting of several self-attention layers and regularization layers. Then, the encoder's output matrix (encoder_output) is processed. L Two pooling operations are performed: one to extract the first bit of the encoded result cls_token, and the other to extract the maximum value of each bit's encoded result bert_output.
[0012] cls_token = encoder_output L [0] (2)
[0013] bert_output=max(encoder_output L ,dim=1) (3)
[0014] This results in two different feature vector representations for the same word. Finally, the feature vectors representing the merchant's behavior, the two text feature vectors representing the product description, and the two text feature vectors representing the chat history are concatenated, and a fully connected neural network is used for non-linear transformation to obtain a comprehensive sample feature vector, output_emb.
[0015] Further, the sample risk identification in step (2) specifically includes: considering the expected model, dividing the samples into two categories, risk-free samples and risky samples, the dimension of the sample feature vector output_emb extracted in step (1) needs to be converted into two dimensions through a fully connected neural network. Then, the Softmax function is used on the two-dimensional vector to calculate the probability that the sample belongs to the risk-free class and the risky class, that is, the model's classification score for the sample. The cross-entropy loss function CE_Loss is used during training:
[0016] L CE (x)=-∑y label log(σ(f θ (x)) (4)
[0017] Among them, y label This is the true label corresponding to the sample. Finally, a manually set threshold is compared with the sample's classification score. If the sample's score in the risk class is greater than the threshold, the sample is determined to have a fraud risk.
[0018] Furthermore, the model iteration based on continuous learning described in step (3) specifically includes two parts of continuous learning:
[0019] In the first part, all new risk point samples are forward-propagated through the online model and the new model respectively, yielding their corresponding comprehensive feature vectors and the model's classification score for each sample. Then, a filter is used to select new samples with scores greater than a manually set threshold in the risk class for knowledge distillation. Knowledge distillation includes alignment operations at the comprehensive feature vector level and the classification score level.
[0020] First, for the combined feature vectors of the new and old samples, we calculate the cosine distance between them and use it to construct the loss function KF_Loss, aiming to make the combined feature vectors extracted by the new model as close as possible to those extracted by the old model. The loss function is:
[0021]
[0022] in, and These are the feature vectors extracted from the new sample in the new model and the old model, respectively.
[0023] Secondly, the classifier output of the model is adjusted by temperature. The output is divided by the temperature T (T is a hyperparameter, which can be set to 0.8), and then the classification score is obtained by using the Softmax function.
[0024]
[0025] The cross-entropy loss function KD_Loss is constructed using the classification scores of the old model as labels and the scores of the new model as predicted values, to adjust the updates of the new model parameters to align with the old model. The specific form of the cross-entropy loss function is:
[0026]
[0027] in and These represent the classification scores obtained by the new sample in the new model and the old model, respectively.
[0028] The second part involves selecting risk samples whose scores in the old risk point data are higher than a threshold (human-set) based on the old model's classification scores, and then randomly sampling from these samples. The sampling results are then displayed as `balck_sample`. old With new risk point data black_sample new The risk-free data white_new is mixed to form a new training set train_set new :
[0029] train_set new =balck_sample' old ∪black_sample new ∪white_new (8)
[0030] Update the model using a new training set so that the model learns directly from historical risk information.
[0031] Finally, the loss function L used during model training is:
[0032] L = L CE +λ1L KF +λ2L KD (9)
[0033] Among them, L CE Including the classification error between new and old data, L KF and L KD Let λ1 and λ2 be the alignment loss of the new risk point data on the old and new models, respectively. KF and L KD The corresponding weights are set manually.
[0034] Based on the above-mentioned e-commerce fraud identification method, the present invention also provides an e-commerce fraud identification system, specifically including a sample feature extraction module, a risk identification module, a knowledge distillation module, and a sample replay module. The sample feature extraction module performs the sample feature extraction operation in step (1); the risk identification module performs the risk identification operation in step (2); and the knowledge distillation module and the sample replay module perform the model iteration operation based on continuous learning in step (3).
[0035] The sample feature extraction module includes several independent BERT models, which extract features from merchant behavior, chat logs, etc.
[0036] The risk identification module includes a fully connected neural network, which is used to determine whether a sample has a fraud risk and to give a corresponding probability score.
[0037] The knowledge distillation module includes the following sub-modules: a module for filtering new risk point data, a module for aligning features extracted from online models and new models, and a module for aligning classification and scoring of online models and new models.
[0038] The sample replay module includes a device for selecting a portion of historical risk point data and mixing it with the latest risk point samples for training, so that the new model can directly learn historical risk information.
[0039] The present invention has at least the following beneficial effects:
[0040] This invention presents a continuous learning-based e-commerce fraud detection method that incorporates a continuous learning framework into the fraud detection model. The combination of merchant behavioral features, product description text features, and chat text data features can fully extract information contained in the transaction process. A temperature regulation mechanism smooths the output scores during model training, increasing the distribution entropy and allowing the model to acquire more information. Knowledge distillation guides the new model parameters during updates, allowing them to refer to the old model's sample processing methods during inference—a means of preserving historical knowledge. Sample replay allows the model to directly access historical data information with limited additional storage resources. These two points improve the accuracy of the new model in identifying historical risk points and effectively mitigate the catastrophic forgetting problem caused by model updates. The continuous learning framework in this invention can be easily adjusted and applied to other practical scenarios, giving the invention good versatility.
[0041] Other advantages, objectives and features of the present invention will be apparent in part from the following description, and in part from the understanding of those skilled in the art through study and practice of the invention. Attached Figure Description
[0042] Figure 1This is a block diagram of the e-commerce fraud identification method based on continuous learning according to the present invention.
[0043] Figure 2 A schematic diagram of sample feature extraction is shown.
[0044] Figure 3 A flowchart of the filter selection process is shown. Detailed Implementation
[0045] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description, and performance testing and analysis of the method of the present invention will be provided.
[0046] like Figure 1 As shown, this embodiment of the invention provides an e-commerce fraud identification method based on continuous learning, which includes four important components: a sample feature extraction module that extracts information related to the transaction from merchant behavior data, product descriptions, and chat records between buyers and sellers; a sample risk identification module that integrates the feature information extracted by the model and detects whether the sample contains fraud risk; a knowledge distillation module that aligns the output results of the screened samples in the old and new models; and a sample replay module that stores historical risk point data according to conditions for the latest training.
[0047] For merchant behavior, it is first converted into a numerical vector format using statistical methods, and then the information is extracted into the wide_output layer through the behavior feature layer. The behavior feature layer consists of two fully connected neural networks connected by a ReLU activation function.
[0048] wide_output=f2(Relu(f1(wide_x))) (1)
[0049] Product descriptions and chat logs between merchants and consumers are text data, processed in two ways. One method is to directly convert the text into word vectors `chat_vec` and `prod_vec`. The other method is to retain the word indices `chat_ids` and `prod_ids` in the text and retrieve the corresponding word vectors from an existing vocabulary. In reality, the length of each sentence varies, so the text data needs to be truncated or padded. Specifically, when a uniform sentence length of 30 is set, for a sentence {w1,...,w... n If n≥30, then retain {w1,...,w}. 30 If n < 30, then the result is padded to {w1, ..., w}. nThe total length is 30, defined by the sequence {,0,…,0}. Positional encoding (position_embeddings) is added to the processed text vector, and then the encoded text information is obtained through an encoder structure containing multiple layers of self-attention mechanisms and regularization. Specifically, the text encoding result h of the l-th layer... t for:
[0050] h l =BertEncoder(f(BertSelfAttention(h l-1 (2)
[0051] Where f is a combination of a series of regularization operations. Then, the result is pooled along the text sequence dimension to obtain cls_token, which serves as the basis for subsequent classification; simultaneously, pooling is performed along the feature dimension to obtain the output bert_output of the text feature extraction module.
[0052] cls_token = encoder_output L [0] (3)
[0053] bert_output=max(encoder_output L ,dim=1) (4)
[0054] Where L is the number of layers in the encoder.
[0055] The feature vectors extracted from merchant behavior, product descriptions, and chat logs are concatenated and then passed through a fully connected neural network to obtain the intermediate layer vector representation output_emb of the sample in the model.
[0056] After obtaining the intermediate layer vector representation of the sample, it is transformed into a two-dimensional vector through a fully connected neural network, corresponding to the number of classification categories. Then, the Softmax function is used to obtain the probability of whether the sample contains fraud risk, which is the classification score.
[0057] score = Softmax(W T output_emb+b) (5)
[0058] During training, the cross-entropy loss function CE_Loss is calculated between the model's output probability scores and the true labels of the samples to update the model parameters. The specific form is as follows:
[0059] L CE (x)=-∑y label log(σ(f θ (x)) (6)
[0060] During model training, knowledge distillation is achieved by introducing two additional loss functions. Sample data is propagated forward through the online model (the old model) to obtain its intermediate layer feature vectors and the online model's classification score for them. To remove noise from the online model, samples with scores greater than a threshold (set to 0.9 here) and a true label of 1 are selected from each batch. These samples contain fraud risk and are identified by the online model, indicating that their information can be correctly reflected by the online model, and the new model should learn this identification process. For the selected samples, the cosine similarity between their intermediate layer feature vectors on the old and new models is calculated, and this is used to construct the loss function KF_Loss:
[0061]
[0062] in, and These are the feature vectors extracted from the new sample in the new model and the old model, respectively. By minimizing the KF_Loss loss function, the intermediate layer output of the new model for the sample is made to converge with the online model, meaning that the new model has acquired some knowledge from the online model.
[0063] After obtaining the classifier output of the new model, the output is first processed using a temperature adjustment mechanism. Specifically, the output vector is divided by the temperature T before calculating the classification score.
[0064]
[0065] Generally, T is set to a real number greater than 1 to reduce the difference between the classifier outputs corresponding to different categories. This increases the weight of negative information and the entropy of the distribution. Temperature regulation is not needed during the inference phase, and the corresponding inference classification scores will be closer to 0 or 1, which is beneficial for later thresholding and the model's final classification result. Furthermore, a loss function KD_Loss is constructed between the classification scores of the filtered samples on the old and new models:
[0066]
[0067] in, and These are the classification scores obtained by the new sample in the new model and the online model, respectively. This function is similar in form to the cross-entropy loss function. By minimizing this loss function, the score distributions of the online model and the new model can be aligned, thereby achieving the effect of obtaining historical risk point information.
[0068] In addition to the new risk point sample data, this example conditionally retains some historical risk point samples for direct training of the new model. During the selection process, since the vast majority of the daily generated data consists of risk-free samples, the proportion of risky samples is extremely small, and the model should focus on risk point features, there is no need to retain historical risk-free samples. For historical risky samples, those that scored below a threshold (set to 0.9 in the experiment) during previous training are first removed, as these samples cannot effectively reflect the correct information of the old model. Then, a certain proportion of the remaining samples are randomly selected to ensure a small number of historical samples are drawn, avoiding excessive storage resources. Finally, the extracted historical risk point samples are mixed with the new dataset to form a new training set, `train_set`. new Jointly train the new model.
[0069] In summary, the loss function used during model training is:
[0070] L = L CE +λ1L KF +λ2L KD (10)
[0071] Wherein, λ1 and λ2 are the weights corresponding to KF_Loss and KD_Loss, respectively, and can be adjusted according to the experimental results.
[0072] The embodiments of this application also provide an e-commerce fraud identification device based on continuous learning, including: a sample feature extraction module, which includes a fully connected neural network for processing behavioral features and a BERT network for processing text information such as chat logs and product descriptions, and the extracted features are concatenated to form a sample feature vector.
[0073] The knowledge distillation module includes a structure for temperature regulation of the classifier output, a filter structure for selecting high-scoring samples, and a loss function that performs alignment operations between the intermediate layer feature vectors of the old and new models and the sample classification scores.
[0074] The sample replay module includes a filter structure for selecting historical risk point samples. This module mixes high-scoring historical risk point samples with new risk point samples to construct a brand-new comprehensive training set and train the new model together.
[0075] The sample risk identification module contains a fully connected neural network, which is used to train the mapping relationship between the spliced features after sample extraction and the data labels, and outputs the information classification results through the Softmax activation function and the set threshold.
[0076] Embodiments of this application also provide a cross-lingual multimodal information fusion apparatus, including:
[0077] Large-scale processors, computing units, and storage servers are used to execute e-commerce fraud detection methods based on continuous learning; large-scale processors and computing units are used for network construction, training, testing, and application; large-scale storage servers are used to store and retrieve the data required by the e-commerce fraud detection methods based on continuous learning.
[0078] To verify the performance of this method on an e-commerce fraud dataset, we selected an e-commerce fraud risk dataset from within Alibaba Group.
[0079] The e-commerce fraud dataset primarily originates from the Xianyu platform and consists mainly of merchant behavior, product descriptions, and chat logs between buyers and sellers. Merchant behavior data has been converted into vector format, while text information corresponds to its respective word vector format, vocabulary index format, and mask. Due to the need for continuous learning, `emb` and `score` fields have been added, representing the intermediate layer feature vector and classification score of new risk point samples in the online model, respectively. A large amount of sample data is generated daily, with risk-free samples accounting for over 98% and risk samples being extremely rare. Therefore, during model training, black and white samples are extracted at a 1:4 ratio, totaling approximately 700,000 samples. In addition, a certain number of historical risk point samples, totaling approximately 750,000 samples, need to be added to the latest training set. The test set also uses real-world data from business scenarios and is divided into two parts. Test data one consists of historical risk samples that the online model can correctly identify, i.e., all black samples, totaling approximately 20,000 samples. The new model is required to have the highest possible recognition rate for these historical risks to ensure the model's retention of historical knowledge. The second test data set consists of all real business data from a single day, including a large number of risk-free samples and a small number of risky samples. The accuracy of the new model on this dataset will be observed.
[0080] To verify the superiority of this method, this embodiment was compared with the following common continuous learning methods on Alibaba Group's e-commerce fraud dataset: LWF (excerpted from "Z.Li and D.Hoiem, "Learning without forgetting," in ECCV. Springer, 2016, pp. 614-629."), MIR (excerpted from "R.Aljundi, E.Belilovsky, T.Tuytelaars, L.Charlin, M.Caccia, M.Lin, L.Page-Caccia, Online continual learning with maximal interfered retrieval, in: Advances in Neural Information Processing Systems 32, 2019, pp. 11849-11860."), DER++ (excerpted from "Buzzega P, Boschini M, Porrello A, et al. Dark experience for general continual learning: a strong, simple baseline[J]. Advances in neural information Processing Systems, 2020, 33: 15920-15930. This embodiment uses recall and precision as evaluation metrics to measure the performance of each algorithm. Recall measures the proportion of historical risk points that the new model can identify, specifically for test data one. Precision examines the accuracy with which the model identifies risks, and its calculation formula is:
[0081]
[0082] This applies to test data two. Since the number of risk-free samples in test data two far exceeds the number of risky samples, even if the new model has a low misidentification rate in risk-free samples, the absolute number of samples misidentified as risky samples is still relatively high compared to the number of risky samples. Therefore, the accuracy should not be too low.
[0083] The results of the model comparison experiment are shown in Table 1. The LWF method mainly learns from historical information through knowledge distillation, but its ability to remember historical risk points is the weakest. MIR and DER++ are both sample replay methods. The former uses temporary model updates to select older samples that have been significantly affected for replay, while the latter optimizes by narrowing the output gap between historical samples and the old and new models. Both methods have recall rates 3.46 and 7.78 percentage points higher than the LWF method, respectively. This indicates that compared to knowledge distillation, sample replay allows the new model to directly access historical samples during training, thus better preserving the memory of historical risk information. The framework proposed in this invention combines knowledge distillation and sample replay methods, and uses filters to select more reasonable samples for training, resulting in a significantly higher recall rate on historical samples than existing continuous learning methods. The recognition accuracy of all four methods reached above 0.5, meeting the requirements of actual business operations.
[0084] Table 2 presents a comparison of ablation experiment results, with a classification threshold of 0.9. Among them, experiment CL... * The results show that the continuous learning method, which incorporates KF_Loss and KD_Loss, improved the recall rate on historical risk points by 0.0223 compared to a completely retrained model, demonstrating the advantages of continuous learning. (Experiment CL) * (temperature) and CL * (replay) represents experiments conducted using either the temperature regulation mechanism or sample replay separately, based on knowledge distillation, and the results demonstrate the effectiveness of both. Finally, the fusion of all continuous learning modules was tested, achieving the highest recall rate on historical risk point data.
[0085] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
[0086] Table 1. Comparison of experimental results for different continuous learning frameworks.
[0087]
[0088] Table 2. Comparison of ablation test results.
[0089]
Claims
1. A method for identifying e-commerce fraud based on continuous learning, characterized in that, The specific steps are as follows: (1) Sample feature extraction; using existing pre-trained vocabulary and text feature extractor, the text information generated during the merchant's transaction process is encoded and concatenated with the merchant's behavioral features to obtain sample features; the pre-trained vocabulary is the set of vector representations corresponding to words; the text feature extractor is selected The encoder module in the model; (2) Sample risk identification; The sample features learned by the model are scored by a binary classifier, and an appropriate threshold is set to obtain the final classification result; (3) Model iteration based on continuous learning; by using knowledge distillation and sample replay methods, the parameters of the new model are made to converge with the parameters of the old model during the training process, so as to retain the ability of the new model to identify historical risks as much as possible while learning the feature information of new risk points; The sample risk identification in step (2) specifically includes: dividing the samples into two categories: risk-free samples and risky samples; and extracting the sample feature vectors from step (1). The dimension is transformed into a two-dimensional vector using a fully connected neural network; then the two-dimensional vector is used... The function calculates the probability that a sample belongs to the risk-free class or the risk class, i.e., the model's classification score for the sample; the cross-entropy loss function is used during training. , (4) in, The true label corresponding to the sample is determined; finally, the set threshold and the sample's classification score are compared. If the sample's score in the risk class is greater than the threshold, the sample is determined to have fraud risk. The model iteration based on continuous learning described in step (3) specifically includes two parts of continuous learning: In the first part, all new risk point samples are forward-propagated through both the old and new models to obtain their corresponding comprehensive feature vectors and the model's classification score for the samples. Then, a filter is used to select new samples with scores greater than a threshold in the risk class for knowledge distillation. Knowledge distillation includes alignment operations at the comprehensive feature vector level and the classification score level. First, for the combined feature vectors of the new and old samples, the cosine distance between them is calculated, and a loss function is constructed based on this distance to make the combined feature vectors extracted by the new model as close as possible to those extracted by the old model; the loss function is: , (5) in, and These are the feature vectors extracted from the new sample in the new model and the old model, respectively. Secondly, the classifier output of the model is adjusted by temperature by dividing the output by the temperature. Then through The function yields the classification score. : (6) The cross-entropy loss function is constructed using the classification scores of the old model as labels and the scores of the new model as predicted values, in order to adjust the updates of the new model parameters to align with the old model; the form of the cross-entropy loss function is: , (7) in, and These are the classification scores obtained by the new sample in the new model and the old model, respectively; The second part involves selecting risk samples whose scores in the old risk point data are higher than a threshold based on the old model's classification scores, and then randomly sampling from these samples; the sampling results are then... With new risk point data Risk-free data Mixing to form a new training set : (8) Update the model using a new training set so that the model can directly learn information about historical risks; Finally, the loss function used during model training for: (9) in, Classification error including both old and new data and The alignment loss of new risk point data on the old and new models, and They are respectively and The corresponding weights are set manually.
2. The e-commerce fraud identification method according to claim 1, characterized in that, The sample feature extraction described in step (1) specifically includes: extracting features from merchant behavior, extracting features from product description text information, and extracting features from chat information between buyers and sellers; finally, concatenating the various feature vectors and performing a nonlinear transformation; wherein: For merchant behavior data Two fully connected neural networks are used. and Activation functions extract behavioral features : (1) For text data such as product descriptions and chat messages between buyers and sellers, text feature vectors can be extracted using one of two methods: one is to use... The model first converts text data into word vectors, then directly retrieves the word vector corresponding to each word from an existing pre-trained vocabulary by using the word index in the text data; finally, it adds positional encoding to the word vectors. The encoding is performed through an encoder, which consists of several self-attention layers and regularization layers; then the output matrix of the encoder is processed. Perform two pooling operations, one for each of the first and second encoded results. And take the maximum value of each bit in the encoding result. : (2) , (3) This results in two different text feature vector representations for the same text data; Finally, the feature vectors of merchant behavior, product description text, and buyer-seller chat information are concatenated and then transformed using a fully connected neural network to obtain a comprehensive sample feature vector. .
3. The e-commerce fraud identification system according to claim 1, characterized in that, Specifically, it includes a sample feature extraction module, a risk identification module, a knowledge distillation module, and a sample replay module; wherein, the sample feature extraction module performs the sample feature extraction operation in step (1); the risk identification module performs the risk identification operation in step (2); and the knowledge distillation module and the sample replay module perform the model iteration operation based on continuous learning in step (3).
4. The e-commerce fraud detection system according to claim 3, characterized in that: The sample feature extraction module includes several independent The model extracts features from merchant behavior, product descriptions, and chat information between buyers and sellers. The risk identification module includes a fully connected neural network, which is used to determine whether a sample has a fraud risk and to give a corresponding probability score; The knowledge distillation module includes the following sub-modules: a module for filtering new risk point data, an alignment module for extracting features from old and new models, and an alignment module for classifying and scoring old and new models. The sample replay module is configured to select a portion of historical risk point data and mix it with the latest risk point samples for training, so that the new model can directly learn historical risk information.
Citation Information
Patent Citations
Additable behavior recognition system and method based on data enhancement and knowledge distillation
CN114638289A