Telecommunication fraud prediction method and system, computer and storage medium

Through combining algorithms and a variational autoencoder that improves the loss function, the telecom fraud data set is processed, the equalization data set is generated and the prediction model is trained, which solves the problem of uneven sample distribution in the telecom fraud data, and realizes effective learning of a few sample types and data balance processing.

CN120069892APending Publication Date: 2025-05-30JIANGXI COLLEGE OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411967194.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology has the problem of uneven sample distribution in the processing of telecom fraud data, which makes the model unable to effectively learn the characteristics of a few samples and then identify a few fraudulent samples.

Method used

The fraud data set is combined and processed through a combination algorithm to generate an equalization data set, and the balanced data set is self-encoded using a variational autoencoder with improved loss function to generate a new fraud data set. Then, the decision tree is added iteratively and the objective function is optimized, and the prediction model is trained to output whether the user has been cheated.

Benefits of technology

By deeply cleaning the data set and increasing the number of samples in a few categories, the data imbalance problem is solved, and by generating new samples similar to the original data, the potential characteristics of the data are effectively learned, and the problem of positive and negative samples is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069892A_ABST
    Figure CN120069892A_ABST
Patent Text Reader

Abstract

The invention provides a telecommunication fraud prediction method and system, a computer and a storage medium, and the method comprises the steps: obtaining a fraud data set, and carrying out the combination processing of the fraud data set through employing a combination algorithm, and obtaining a balanced data set; performing self-encoding on the balanced data set by using a variational self-encoder of an improved loss function to generate a new fraud data set similar to the balanced data set; iteratively adding decision trees based on the new fraud data set, and optimizing a target function to train each decision tree to obtain a prediction model; and obtaining the feature data of the transaction of the user, inputting the feature data into the prediction model, and outputting whether the user has a fraud behavior, thereby solving the problem of data imbalance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method, a system, a computer, and a storage medium for predicting telecommunications fraud. Background Art

[0002] In recent years, telecommunications network fraud has continued to be highly prevalent, and "telecommunications network fraud" has become one of the main factors that make people feel "not very safe" or "unsafe". Telecommunications fraud criminals purchase or produce pseudo base stations and hacker software at low prices, and implement fraud on a large number of potential victims, obtaining huge amounts of criminal proceeds in a short time.

[0003] In terms of predicting the risk of telecommunications fraud, currently mainly machine learning-based methods are used, and these methods can be divided into two major categories: supervised learning and unsupervised learning. Unsupervised learning methods usually use the K-means clustering algorithm to identify possible fraud patterns.

[0004] Currently, the main approaches to combating telecommunications fraud crimes are explored from the perspectives of information flow, fund flow, evidence chain, etc. However, there is an extremely unbalanced sample distribution in telecommunications fraud data, resulting in the model being unable to effectively learn the characteristics of minority class samples (data of defrauded users), so that the model cannot identify the minority of defrauded samples from a large number of samples. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method, a system, a computer, and a storage medium for predicting telecommunications fraud to solve the technical problems in the prior art.

[0006] On the one hand, the invention provides the following technical solution. A method for predicting telecommunications fraud, the method comprising:

[0007] Obtain a fraud data set, and perform combined processing on the fraud data set by using a combined algorithm to obtain a balanced data set;

[0008] Perform autoencoding on the balanced data set by using a variational autoencoder with an improved loss function to generate a new fraud data set similar to the balanced data set;

[0009] Iteratively add decision trees based on the new fraud data set, and optimize the objective function to train each of the decision trees to obtain a prediction model;

[0010] Obtain the characteristic data of a user's transaction, input the characteristic data into the prediction model, and output whether the user has been defrauded.

[0011] Compared with the prior art, the beneficial effects of the present application are as follows: By using a combination algorithm to perform combination processing on the fraud data set to obtain a balanced data set, the fraud data set is deeply cleaned, thereby solving the problem of data imbalance. It not only increases the number of minority-class samples but also improves the problem that the minority-class samples introduced by the SMOTE algorithm are repeated with the majority-class samples. By using a variational autoencoder with an improved loss function to perform autoencoding on the balanced data set to generate a new fraud data set similar to the balanced data set, it can not only effectively learn the latent features of the data but also generate new samples similar to the original data, solving the problem of imbalance between positive and negative samples in the data set.

[0012] Further, the step of using the combination algorithm to perform combination processing on the fraud data set includes:

[0013] Using the Synthetic Minority Over-sampling Technique (SMOTE) to oversample the minority-class samples of the fraud data set to generate a new minority-class sample set;

[0014] Using the Edited Nearest Neighbor (ENN) algorithm to clean the minority-class sample set to obtain the judgment category of the minority-class sample set;

[0015] Judge whether the type of the minority-class sample set is the same as the judgment category. If the type of the minority-class sample set is the same as the judgment category, output the minority-class sample set.

[0016] Further, the step of using the Edited Nearest Neighbor (ENN) algorithm to clean the minority-class sample set includes:

[0017] Find multiple category samples at a preset distance in the feature space for each sample in the minority-class sample set, and count the number of each category among the multiple category samples;

[0018] If the majority-class samples among the multiple category samples are in the majority, the predicted category is the majority class;

[0019] If the minority-class samples among the multiple category samples are in the majority, the predicted category is the minority class.

[0020] Further, the representation of the minority-class sample set is:

[0021] X S = X iid +(X iid - X icd )×rand(0, 1)

[0022] where X S is the minority-class sample set, X iid is an arbitrarily selected sample, X icd is one of the current traversed nearest neighbor points, Xiid a nearest neighbor points in the fraud dataset are X inn , X icd ∈ X inn , a is the number of samples in the preset distance category, and rand(0, 1) returns a vector.

[0023] Furthermore, the step of auto - encoding the balanced dataset using the variational auto - encoder with the improved loss function includes:[[]]

[0024] Using the encoder of the variational auto - encoder to map the balanced dataset to the latent space and generate latent variables;

[0025] Calculating a new fraud dataset based on the latent variables using the differential operation formula.

[0026] Furthermore, the improved loss function of the variational auto - encoder includes:[[]]

[0027] L(x) = -FLoos(P(X′ v ∣X′)) + D KL (P(X′ v ∣X′)||Q(X′ v ∣X′))

[0028] where L(x) is the improved variational auto - encoder loss function, -FLoos(P(X′ v ∣X′)) is the loss function modified by binary cross - entropy, and D KL (P(X′ v ∣X′)||Q(X′ v ∣X′)) is the KL divergence between the true posterior distribution and the approximate posterior distribution;

[0029] The differential operation formula includes:[[]]

[0030]

[0031] where X v ′ is the new fraud dataset, μ and δ are the mean and standard deviation of the latent variables, is the noise sampled from the standard normal distribution.

[0032] Furthermore, the expression of the prediction model is:[[]]

[0033]

[0034] where is the prediction model, k represents the number of decision trees, F represents the set of all decision trees, and f k is the k - th decision tree generated in the k - th iteration, and X iFeature vectors represented as a new fraud dataset;

[0035] The objective function of the prediction model is:

[0036]

[0037] Where is the objective function, Y i is the corresponding label of the new fraud dataset, ∑ n Ω(f n ) is the regularization penalty function, and l is the error function.

[0038] In a second aspect, the invention provides the following technical solution. A telecom fraud prediction system, the system includes:

[0039] A processing module, configured to obtain a fraud dataset, and perform combined processing on the fraud dataset by using a combined algorithm to obtain a balanced dataset;

[0040] A generation module, configured to perform autoencoding on the balanced dataset by using a variational autoencoder with an improved loss function to generate a new fraud dataset similar to the balanced dataset;

[0041] A training module, configured to iteratively add decision trees based on the new fraud dataset and optimize the objective function to train each of the decision trees to obtain a prediction model;

[0042] A judgment module, configured to obtain feature data of a user's transaction, input the feature data into the prediction model, and output whether the user has been defrauded.

[0043] In a third aspect, the invention provides the following technical solution. A computer includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned telecom fraud prediction method is implemented.

[0044] In a fourth aspect, the invention provides the following technical solution. A storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned telecom fraud prediction method is implemented. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 Flow chart of the telecommunications fraud prediction method provided by the first embodiment of the present invention;

[0047] Figure 2 Block diagram of the structure of the telecommunications fraud prediction system provided by the second embodiment of the present invention;

[0048] Figure 3 Schematic diagram of the hardware structure of a computer provided by the third embodiment of the present invention.

[0049] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings. Detailed implementation manners

[0050] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals throughout are the same or similar elements or elements having the same or similar functions. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the embodiments of the present invention and should not be construed as limiting the present invention.

[0051] In the description of the embodiments of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention.

[0052] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality" means two or more unless otherwise specifically defined.

[0053] Embodiment 1

[0054] In the first embodiment of the present invention, please refer to Figure 1 As shown, a telecommunications fraud prediction method includes the following steps S01 to S04:

[0055] S01. Obtain a fraud data set, and perform combination processing on the fraud data set by using a combination algorithm to obtain a balanced data set;

[0056] Specifically, the step of performing combination processing on the fraud data set by using a combination algorithm includes:

[0057] S11, oversample the minority-class samples in the fraud data set using the Synthetic Minority Over-sampling Technique (SMOTE) to generate a new minority-class sample set;

[0058] S12, clean the minority-class sample set using the Edited Nearest Neighbor (ENN) algorithm to obtain the judgment categories of the minority-class sample set;

[0059] S13, determine whether the type of the minority-class sample set is the same as the judgment category. If the type of the minority-class sample set is the same as the judgment category, output the minority-class sample set.

[0060] More specifically, the step of cleaning the minority-class sample set using the Edited Nearest Neighbor (ENN) algorithm includes:

[0061] S121, find multiple category samples at a preset distance in the feature space for each sample in the minority-class sample set, and count the number of each category among the multiple category samples;

[0062] S122, if the majority-class samples among the multiple category samples are in the majority, the predicted category is the majority class;

[0063] S123, if the minority-class samples among the multiple category samples are in the majority, the predicted category is the minority class. Specifically, the minority-class sample set is represented as:

[0064] X S = X iid +(X iid - X icd )×rand(0, 1)

[0065] where X S is the minority-class sample set, X iid is an arbitrarily selected sample, X icd is one of the current traversed nearest neighbor points, a nearest neighbor points of X iid in the fraud data set are X inn , X icd ∈X inn , a is the number of category samples at the preset distance, and rand(0, 1) returns a vector.

[0066] In this embodiment, the fraud data set (X) includes: the distance between the bank card transaction location and home, with the unit of hm; the distance from the last transaction, with the unit of hm; the ratio of the most recent transaction to the median price of previous transactions, with the value range between 0 and 1; whether the transaction occurs at the same merchant, where 0 means "no" and 1 means "yes"; whether the transaction is made through a chip (bank card), where 0 means "no" and 1 means "yes"; whether a PIN code is used during the transaction, where 0 means "no" and 1 means "yes"; whether it is an online transaction order, where 0 means "no" and 1 means "yes".

[0067] Among them, the Synthetic Minority Oversampling Technique is SMOTE, and the Edited Nearest Neighbours algorithm is ENN.

[0068] In specific applications, SMOTE (Synthetic Minority Oversampling Technique) and ENN (Edited Nearest Neighbours) are used to perform combined sampling on the data set. Specifically, the combined algorithm is SMOTE-ENN. Based on SMOTE, the ENN algorithm is used to deeply clean the data generated by the SMOTE method, and finally a balanced data set is obtained. More specifically:

[0069] Assume that the fraud data set is X, and the minority class samples X S generated by SMOTE are:[[]]

[0070] X S = X iid +(X iid - X icd )×rand(0, 1)

[0071] Among them, X S is the minority class sample set, X iid is an arbitrarily selected sample, X icd is one of the current traversed nearest neighbor points, and the a nearest neighbor points of X iid in the fraud data set are X inn , X icd ∈ X inn , a is the number of preset distance category samples, rand(0, 1) returns a vector, and the elements in this vector all follow a random value of the 0-1 distribution. Then, the ENN algorithm is used to clean the generated minority class samples, and it is judged whether X S is the same as the prediction type of the a-nearest neighbor algorithm, that is, find X SFor the a samples closest in distance, count the number of each category among these a samples. If the majority class samples among the a nearest neighbor samples are in the majority, then predict that the category of X S is the majority class. If the minority class samples among the a nearest neighbor samples are in the majority, then predict that the category of X S is the minority class. Then compare whether the predicted type of X S is the same as that predicted by the a-nearest neighbor algorithm. If the predicted category is the same as the original minority class label of X S , retain the sample. If the predicted category is different from the original minority class label of X S , delete the sample. Finally, output the new dataset (balanced dataset X′).

[0072] It should be noted that the SMOTE-ENN algorithm is used for combined sampling of the credit card dataset. This method further uses the ENN algorithm to deeply clean the data generated by SMOTE on the basis of SMOTE oversampling, thus solving the problem of data imbalance. It not only increases the number of minority class samples but also improves the problem of duplication of minority class samples and majority class samples introduced by the SMOTE algorithm.

[0073] S02. Use a variational autoencoder with an improved loss function to perform autoencoding on the balanced dataset to generate a new fraud dataset similar to the balanced dataset;

[0074] Specifically, the steps of using a variational autoencoder with an improved loss function to perform autoencoding on the balanced dataset include:

[0075] Use the encoder of the variational autoencoder to map the balanced dataset to the latent space and generate latent variables;

[0076] Calculate the new fraud dataset based on the latent variables using the differential operation formula.

[0077] Specifically, the improved loss function of the variational autoencoder includes:

[0078] L(x) = -FLoos(P(X′ v ∣X′)) + D KL (P(X′ v ∣X′)PQ(X′ v ∣X′))

[0079] where L(x) is the improved loss function of the variational autoencoder, -FLoos(P(X′ v ∣X′)) is the loss function modified from binary cross-entropy, and D KL (P(X′ v ∣X′)PQ(X′ v ∣X′)) is the KL divergence between the true posterior distribution and the approximate posterior distribution.

[0080] The differential operation formula includes:

[0081]

[0082] where X v ′ is the new fraud dataset, μ and δ are the mean and standard deviation of the latent variables, and is the noise sampled from the standard normal distribution.

[0083] In this embodiment, the Variational Autoencoder (VAE) is remarkable in generating new samples, learning latent features, improving model robustness, etc. It can not only effectively learn the latent features of the data, but also generate new samples similar to the original data. The VAE with the improved loss function is used to auto - encode the balanced dataset X′, that is, given the observed data X′, find the distribution X v ′ of the latent variables. The basic structure of the VAE includes an encoder, a latent space, and a decoder. First, the encoder maps the input data to the latent space and generates the distribution parameters (mean and variance) of the latent variables, rather than a single latent vector. In this way, the model can capture the complexity of the data. In the latent space, the latent variables are generated by random sampling, which introduces randomness and enables the model to have the generation ability.

[0084] In this embodiment, in order to calculate the posterior distribution, P(X v ′|X′), the VAE uses an inference network to approximate the true posterior distribution Q(X v ′|X′). Since directly calculating the posterior distribution is usually infeasible, the VAE minimizes the difference between the true posterior distribution and the approximate posterior distribution through variational inference, using the KL - divergence as a metric:

[0085]

[0086] To maximize the marginal likelihood P(X′), the variational lower bound is introduced:

[0087] logP(X′)≥FLoos(P(X′ v |X′)) - D KL (P(X′ v |X′)||P(X′))

[0088] The first term is the loss function modified from binary cross - entropy:

[0089]

[0090] where α is the class weight, - log(P(X′ v∣X′)) is the original cross-entropy loss function, (1 - P(X′ v ∣X′)) γ is a parameter for adjusting easy or difficult classification samples. By carefully setting the key value range of α, the weight distribution of positive and negative samples in the model can be effectively regulated. If the value of α is smaller, then the weight for negative samples should be reduced. α can only adjust the weights of positive and negative samples and cannot adjust the weights of easy and difficult classification samples, while adjusting γ can reduce the weight of easy classification samples. y represents the true label of the sample. This loss function solves the problem of imbalance between positive and negative samples in the dataset by assigning different weights to positive and negative samples, and can more accurately measure the similarity between the generated samples and the true samples for an imbalanced dataset. The goal of VAE is to maximize the variational lower bound, which can be rewritten as minimizing the loss function:

[0091] L(x) = -Floos(P(X′ v ∣X′)) + D KL (P(X′ v ∣X′)||Q(X′ v ∣X′))

[0092] To make the model trainable, the reparameterization trick is introduced to convert the sampling process of latent variables into a differentiable operation:

[0093]

[0094] where, X v ′ is the new fraud dataset, μ and δ are the mean and standard deviation of the latent variables, is the noise sampled from the standard normal distribution. The generated new dataset X v ′ is used as the input data for XGBoost (prediction model).

[0095] S03, iteratively add decision trees based on the new fraud dataset, and optimize the objective function to train each of the decision trees to obtain a prediction model;

[0096] Specifically, the expression of the prediction model is:

[0097]

[0098] where, is the prediction model, k represents the number of decision trees, F represents the set of all decision trees, f k is the k-th decision tree generated in the k-th iteration, and X i represents the feature vector of the new fraud dataset.

[0099] The objective function of the prediction model is:

[0100]

[0101] Among them, is the objective function, and Y i is the corresponding label of the new fraud dataset, and ∑ n Ω(f n ) is the regularization penalty function, and l is the error function.

[0102] In this embodiment, the XGBoost algorithm trains multiple decision tree models iteratively and gradually improves the prediction performance of the model by optimizing the objective function. In each round of iteration, XGBoost constructs a new decision tree model based on the residuals between the prediction results of the previous round of model and the true labels. Then, the prediction results of the new model are weighted and added to the results of the previous round of model to obtain the final integrated prediction result. If the dataset containing feature a and with a capacity of b is represented as:

[0103]

[0104] Among them, D is the new dataset X v ′, the dataset has a columns and b rows, and X i represents the features of the new fraud dataset, and Y i is the label of the new fraud dataset, represents the real number field, represents an a-dimensional vector in the real number field.

[0105] The telecommunications fraud risk prediction model based on XGBoost can be expressed as:

[0106]

[0107] Among them, is the prediction model, k represents the number of decision trees, F represents the set of all decision trees, and f k is the kth decision tree generated in the kth iteration, and X i represents the features of the new fraud dataset. The objective function for model optimization is:

[0108]

[0109] Among them, is the objective function, and Y i is the corresponding label of the new fraud dataset, and ∑ n Ω(f n ) is the regularization penalty function, and l is the error function.

[0110] S04, obtain the feature data of the user's transaction, input the feature data into the prediction model, and output whether the user has the behavior of being defrauded.

[0111] In this embodiment, based on the characteristic data of the user's transactions, it is predicted whether the user has been defrauded.

[0112] In summary, a method for predicting telecom fraud has the following effects:

[0113] By using the combination algorithm to perform combination processing on the fraud data set to obtain the balanced data set, the fraud data set is deeply cleaned, thereby solving the problem of data imbalance. It not only increases the number of minority class samples but also improves the problem of repetition of minority class samples and majority class samples introduced by the SMOTE algorithm. By using the variational autoencoder with an improved loss function to perform autoencoding on the balanced data set to generate a new fraud data set similar to the balanced data set, it can not only effectively learn the latent features of the data but also generate new samples similar to the original data, solving the problem of imbalance between positive and negative samples in the data set.

[0114] Embodiment Two

[0115] As Figure 2 shown, in the second embodiment of the present invention, a telecom fraud prediction system is provided. The system includes:

[0116] A processing module 10, configured to obtain a fraud data set and use a combination algorithm to perform combination processing on the fraud data set to obtain a balanced data set;

[0117] A generating module 20, configured to use a variational autoencoder with an improved loss function to perform autoencoding on the balanced data set to generate a new fraud data set similar to the balanced data set;

[0118] A training module 30, configured to iteratively add decision trees based on the new fraud data set and optimize the objective function to train each of the decision trees to obtain a prediction model;

[0119] A judging module 40, configured to obtain the characteristic data of the user's transactions, input the characteristic data into the prediction model, and output whether the user has been defrauded.

[0120] For the telecom fraud prediction system provided by the embodiment of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the system embodiment, reference may be made to the corresponding content in the foregoing method embodiment.

[0121] Embodiment Three

[0122] As Figure 3As shown, in the third embodiment of the present invention, the present invention embodiment provides the following technical solution. A computer includes a memory 202, a processor 201, and a computer program stored on the memory 202 and executable on the processor 201. When the processor 201 executes the computer program, the above-mentioned telecommunications fraud prediction method is implemented.

[0123] Specifically, the above-mentioned processor 201 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0124] Among them, the memory 202 may include a mass storage for data or instructions. By way of example and not limitation, the memory 202 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In appropriate cases, the memory 202 may include removable or non-removable (or fixed) media. In appropriate cases, the memory 202 may be internal or external to the data processing device. In a particular embodiment, the memory 202 is a non-volatile memory. In a particular embodiment, the memory 202 includes a read-only memory (ROM) and a random access memory (RAM). In appropriate cases, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In appropriate cases, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0125] The memory 202 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 201.

[0126] The processor 201 reads and executes the computer program instructions stored in the memory 202 to implement the above-mentioned telecommunications fraud prediction method.

[0127] In some of these embodiments, the computer may further include a communication interface 203 and a bus 200. Among them, as Figure 3 shown, the processor 201, the memory 202, and the communication interface 203 are connected through the bus 200 and complete communication with each other.

[0128] The communication interface 203 is used to implement communication between each module, device, unit, and / or equipment in the embodiments of the present application. The communication interface 203 can also implement data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.

[0129] Bus 200 includes hardware, software, or both, and couples components of a computer together. Bus 200 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 200 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In a suitable case, Bus 200 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0130] Embodiment Four

[0131] In the fourth embodiment of the present invention, in combination with the above-mentioned telecommunications fraud prediction method, the embodiments of the present invention provide the following technical solution: a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned telecommunications fraud prediction method is implemented.

[0132] Those skilled in the art can understand that in the flowchart, the data and / or steps described as or in other ways herein, for example, can be considered as a predefined sequence data table of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0133] More specific examples (non-exhaustive list) of the readable medium include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0134] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0135] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0136] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A telecommunications fraud prediction method, characterized in that: The method comprises: Obtain a fraud data set, and combine the fraud data set using a combination algorithm to obtain a balanced data set; Using a variational autoencoder with an improved loss function to autoencode the balanced data set to generate a new fraud data set similar to the balanced data set; Iteratively adding decision trees based on the new fraud data set, and optimizing the objective function to train each decision tree to obtain a prediction model; Obtain feature data of transactions conducted by the user, input the feature data into the prediction model, and output whether the user has been defrauded.

2. The telecommunications fraud prediction method according to claim 1, characterized in that: The step of combining the fraud data set using a combination algorithm comprises: Oversampling the minority class samples of the fraud data set using a synthetic minority class oversampling technique to generate a new minority class sample set; Using the edited nearest neighbor algorithm to clean the minority class sample set, to obtain the judgment category of the minority class sample set; It is determined whether the type of the minority class sample set is the same as the determination category. If the type of the minority class sample set is the same as the determination category, the minority class sample set is output.

3. The telecommunications fraud prediction method according to claim 2, characterized in that: The step of cleaning the minority class sample set by using the edited nearest neighbor algorithm comprises: Find multiple category samples with a preset distance in the feature space for each sample in the minority class sample set, and count the number of each category in the multiple category samples; If the majority class samples account for the majority among the multiple class samples, the predicted class is the majority class; If the minority class samples account for the majority among the multiple class samples, the predicted class is the minority class.

4. The telecommunications fraud prediction method according to claim 2, characterized in that: The minority class sample set is represented as: X S =X iid +(X iid -X icd )×rand(0,1) Among them, X S is the minority class sample set, X iid For any selected sample, X icd is one of the neighboring points currently traversed, X iid The a nearest neighbor point in the fraud dataset is X inn , X icd ∈X inn , a is the number of preset distance category samples, and rand(0, 1) returns a vector.

5. The telecommunications fraud prediction method according to claim 1, characterized in that: The step of using the variational autoencoder with an improved loss function to autoencode the balanced data set comprises: Mapping the balanced data set to a latent space using an encoder of the variational autoencoder and generating latent variables; A new fraud data set is calculated based on the latent variables using a differential operation formula.

6. The telecommunications fraud prediction method according to claim 5, characterized in that: The improved loss function of the variational autoencoder includes: L(x)=-FLoos(P(X′ v ∣X′))+D KL (P(X′ v ∣X′)||Q(X′ v ∣X′)) Where L(x) is the improved variational autoencoder loss function, -Floos(P(X′ v |X′)) is the loss function modified by binary cross entropy, D KL (P(X′ v |X′)||Q(X′ v |X′)) is the KL divergence between the true posterior distribution and the approximate posterior distribution; The differential operation formula includes: Among them, X v ′ is the new fraud dataset, μ and δ are the mean and standard deviation of the latent variables, is the noise sampled from a standard normal distribution.

7. The telecommunications fraud prediction method according to claim 1, characterized in that: The expression of the prediction model is: in, is the prediction model, k represents the number of decision trees, F represents the set of all decision trees, and f k is the kth decision tree generated in the kth iteration, X i Represented as the feature vector of the new fraud dataset; The objective function of the prediction model is: in, is the objective function, Y i is the corresponding label of the new fraud dataset, ∑ n Ω(f n ) is the regular penalty function, and l is the error function.

8. A telecommunications fraud prediction system, characterized in that: The system comprises: A processing module, used for obtaining a fraud data set, and combining and processing the fraud data set using a combination algorithm to obtain a balanced data set; A generating module, used for performing autoencoding on the balanced data set by using a variational autoencoder with an improved loss function to generate a new fraud data set similar to the balanced data set; A training module, used for iteratively adding decision trees based on the new fraud data set, and optimizing the objective function to train each decision tree to obtain a prediction model; The judgment module is used to obtain feature data of transactions conducted by users, input the feature data into the prediction model, and output whether the user has been defrauded.

9. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the telecommunications fraud prediction method as described in any one of claims 1 to 7 is implemented.

10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the telecommunications fraud prediction method as described in any one of claims 1 to 7.