Telecommunication fraud detection method and device, electronic equipment and storage medium
By updating the reward function of the training model in telecom fraud detection, the problem of insufficient model adaptability is solved, achieving more efficient accuracy and flexibility in telecom fraud detection, and adapting to the rapid changes in fraud methods.
Patent Information
- Application Number
- CN202511193620.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-12-12
AI Technical Summary
Existing machine learning models struggle to adapt quickly to changes in fraud tactics in telecom fraud detection, leading to decreased detection performance and requiring extensive retraining with labeled data.
When the accuracy of the detection model is less than the preset accuracy, the second model is updated and trained based on the first training data to obtain the first reward function. The entropy optimization method is used to learn the reward function consistent with fraudulent behavior, and the first model is updated and trained to reduce the dependence on a large amount of labeled data.
It improves the model's adaptability and detection accuracy, reduces the difficulty of manually designing complex reward functions, and significantly enhances the accuracy and flexibility of telecom fraud detection.
Smart Images

Figure CN121125901A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication information security, and in particular to a telecommunications fraud detection method and device, an electronic device and a storage medium. BACKGROUND
[0002] In recent years, telecommunications network fraud has become increasingly serious, seriously threatening the safety and happiness of the public and becoming a major social problem.
[0003] Currently, telecommunications network fraud detection can be performed based on a machine learning model to identify fraudulent behavior. However, the current machine learning model can learn fraud patterns from historical data, but when fraud methods change and new fraud patterns appear, the model needs to be retrained with a large amount of labeled data to adapt to new fraud methods, which is often difficult to achieve in practical applications, resulting in a decline in the detection performance of the model. Therefore, it is very important to more quickly and accurately identify fraudulent behavior. SUMMARY
[0004] The present application provides a telecommunications fraud detection method, device, electronic device and storage medium to significantly improve the accuracy of telecommunications fraud detection.
[0005] According to an aspect of the present application, a telecommunications fraud detection method is provided, the method comprising:
[0006] If the accuracy of the first model is less than the preset accuracy, the second model is updated and trained based on the first training data to obtain a first reward function; the first model is used to detect whether the call behavior has fraudulent behavior; the first training data includes call behavior trajectories labeled with normal behavior trajectories and fraudulent behavior trajectories of all fraud types; the second model is used to learn a reward function consistent with fraudulent behavior based on the fraudulent behavior trajectory;
[0007] The first model is updated and trained based on the first reward function and the first training data to obtain an updated first model;
[0008] The call behavior is identified based on the updated first model, and whether the call behavior has fraudulent behavior is output.
[0009] According to another aspect of the present application, a telecommunications fraud detection device is provided, the device comprising:
[0010] a function determining module configured to, if it is detected that the accuracy of the first model is less than a preset accuracy, update and train the second model based on first training data to obtain a first reward function; the first model is configured to detect whether a call behavior exists fraud behavior; the first training data comprises a call behavior trajectory labeled with a normal behavior trajectory and a fraud behavior trajectory of all fraud types; the second model is configured to learn a reward function consistent with the fraud behavior based on the fraud behavior trajectory;
[0011] an updating module configured to update and train the first model based on the first reward function and the first training data to obtain an updated first model;
[0012] a detecting module configured to identify the call behavior based on the updated first model and output whether the call behavior exists fraud behavior.
[0013] According to another aspect of the present application, an electronic device is provided, which comprises:
[0014] at least one processor; and
[0015] a memory connected to the at least one processor in communication; wherein,
[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for detecting telecom fraud according to any one of the embodiments of the present application.
[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the method for detecting telecom fraud according to any one of the embodiments of the present application when executed by the processor.
[0018] The technical solution of the embodiment of the application is that if the accuracy of the first model is less than the preset accuracy, the second model is updated and trained based on the first training data to obtain a first reward function; the first model is used to detect whether a call behavior exists fraud; the first training data includes a call behavior trajectory labeled with normal behavior trajectories and fraud behavior trajectories of all fraud types; the second model is used to learn a reward function consistent with fraud behavior based on the fraud behavior trajectory; the automatic acquisition of the first reward function avoids the difficulty of manually designing a complex reward function, improves the adaptability of the model, and reduces the dependence on a large amount of labeled data. Further, the first model is updated and trained based on the first reward function and the first training data to obtain an updated first model. The introduction of the first reward function of the application enables the first model to more flexibly cope with different call behaviors and analysis requirements, so as to identify the call behavior based on the updated first model and output whether the call behavior exists fraud, thereby significantly improving the accuracy of telecom fraud detection.
[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0021] Figure 1 is a flowchart of a telecom fraud detection method according to an embodiment of the application;
[0022] Figure 2 is a structural schematic diagram of a telecom fraud detection device according to an embodiment of the application;
[0023] Figure 3 is a structural schematic diagram of an electronic device for implementing the telecom fraud detection method according to an embodiment of the application. DETAILED DESCRIPTION
[0024] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application, and obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the present application.
[0025] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] Embodiment one
[0027] Figure 1 A flowchart of a telecommunications fraud detection method provided by the embodiments of the present application, the present embodiment can be applicable to the case of telecommunications fraud detection, and the method can be executed by a telecommunications fraud detection device, which can be realized in the form of hardware and / or software, and can be configured in any electronic device with network communication function. As shown in the figure, the method comprises: Figure 1
[0028] S110, if it is detected that the accuracy of the first model is less than the preset accuracy, the second model is updated and trained based on the first training data to obtain a first reward function; the first model is used to detect whether the call behavior exists fraud behavior; the first training data includes call behavior trajectory labeled with normal behavior trajectory and all fraud behavior trajectories of fraud types; the second model is used to learn the reward function consistent with the fraud behavior based on the fraud behavior trajectory.
[0029] The first model can be a graph neural network (GNN) model.
[0030] Specifically, call data of a telecommunication network call environment is acquired, invalid, erroneous or incomplete records are removed by cleaning the call data, feature data of the cleaned call data is extracted, the feature data of the call data includes but is not limited to call objects, basic information, behavior characteristics, call frequency and call duration of a call behavior, a graph data structure of the telecommunication network call environment is constructed based on the extracted feature data of the call data, and the graph data structure of the telecommunication network call environment is taken as first training data. That is, the representation form of the first training data can be the graph data structure of the telecommunication network call environment, the graph data structure is composed of nodes and edges, the nodes are two call objects of a call behavior, and the two call objects of the call behavior are connected to become edges; the attributes of the nodes include at least basic information and behavior characteristics of the call behavior; the attributes of the edges include at least call frequency, call duration and weight of the edges; and the weight of the edges is determined based on the call frequency and the call duration.
[0031] In the process of detecting fraud behavior by using the first model, whether the accuracy of the first model is less than a preset accuracy is detected in real time, the preset accuracy is a minimum accuracy reflecting the accuracy of the detection accuracy of the first model, when the accuracy is less than the preset accuracy, it means that the detection result of the first model is inaccurate, when the accuracy is greater than or equal to the preset accuracy, it means that the detection result of the first model is still accurate, therefore, if the accuracy of the first model is greater than the preset accuracy, the first model does not need to be updated, and the current first model can be continuously used for fraud behavior detection in the future; if it is detected that the accuracy of the first model is less than the preset accuracy, the second model needs to be updated and trained based on the first training data to obtain the first reward function, so that the first model can be accurately updated in the future by using the first reward function.
[0032] In the embodiment, the second model includes a target function, the target function is used to describe the probability of a reward function consistent with the fraud behavior, the second model is updated and trained based on the first training data to obtain the first reward function, including: adjusting the target function based on the first training data, when the target value of the target function is greater than a preset value, determining the reward function corresponding to the target function as the first reward function.
[0033] Specifically, based on the first training data, the target function is solved by using an entropy optimization method, when the target value of the target function is less than the preset value, the parameters of the target function are adjusted according to the target value until the target value of the target function is greater than the preset value, and the reward function corresponding to the target function is determined as the first reward function. The parameters of the target function are also the parameters of the reward function.
[0034] Further, the target function can be expressed by the following formula:
[0035] L(θ)=∑ τ∈DlogP(τ|θ)-logZ(θ);
[0036] Wherein, L(θ) is a target value;D is a fraud behavior trajectory in the first training data;P(τ|θ) is the probability of the fraud behavior trajectory τ in the first training data under the parameter θ;Z(θ) is a normalization constant, and the normalization constant Z(θ) is expressed as: Z(g)=∑ τ∈Ω P(τ|θ), and Ω is the set of all fraud behavior trajectories.
[0037] Meanwhile, the entropy optimization method also ensures the rationality of the reward function through feature constraints, that is, the target function is constrained by a preset constraint condition, and the preset constraint condition is used to describe that the first expected value is the same as the second expected value;The first expected value can be the expected value of the feature function on the fraud behavior trajectory generated by the second model;The second expected value can be the expected value of the feature function on the fraud behavior trajectory in the first training data;The feature function is used to describe the state and action of the fraud behavior;When the target value of the target function is greater than the preset value, the first reward function is constructed by using the feature function corresponding to the target function with the target value greater than the preset value.
[0038] Wherein, the preset constraint condition can be expressed by the following formula:
[0039]
[0040] Wherein, g(s,a) is a feature function of state s and action a; is the first expected value; is the second expected value;By making the two expected values equal, the second model can learn the reward function consistent with the expert behavior. The feature function is gradually adjusted in the training process of the second model, until the target function is maximized, and the preset constraint condition is satisfied.
[0041] In the embodiment of the application, the update of the second model can be understood as a reward function recovery mechanism based on entropy optimization. This mechanism not only improves the adaptability of the model to different fraud scenarios, but also reduces the demand for a large amount of labeled data. In the traditional fraud detection method, manual design of the reward function often requires a large amount of labeled data to verify its effectiveness, and the present application directly learns the reward function through expert-labeled cases, avoiding this problem. At the same time, the reward function optimized by the entropy optimization method can significantly reduce the computational complexity without sacrificing the detection accuracy, thereby providing an efficient and adaptable solution for telecom network fraud detection.
[0042] S120, update and train the first model based on the first reward function and the first training data, and obtain an updated first model.
[0043] Specifically, the first reward function is introduced to update and train the first model, greatly reducing the influence of noise nodes, thereby improving the accuracy of detection.
[0044] Further, based on the first reward function and the first training data, the first model is updated and trained to obtain an updated first model, which can include steps A1-A2:
[0045] Step A1, based on the classification loss function and the reward signal of the first reward function, a target loss function of the first model is constructed.
[0046] Wherein, the classification loss function can be cross-entropy loss function. The target loss function can be expressed as follows:
[0047] L=L1-λf;
[0048] Wherein, L1 is the classification loss function; f is the reward signal of the first reward function; λ is a regularization coefficient, used to balance the influence between the classification loss and the reward signal.
[0049] Step A2, based on the first training data and the target loss function, the parameters of the first model are adjusted until the loss value of the target loss function is less than the preset loss value, and the updated first model is obtained.
[0050] S130, based on the updated first model, the call behavior is identified, and whether the call behavior exists fraud is output.
[0051] Specifically, in the actual application process, the first simulation is used to identify the feature information of the call behavior, so as to accurately output the identification result, and determine whether the call behavior is a fraud behavior through the identification result, so as to avoid unnecessary loss caused by fraud behavior.
[0052] The technical scheme of the embodiment of the application, if the accuracy of the first model is less than the preset accuracy, the second model is updated and trained based on the first training data to obtain a first reward function; the first model is used to detect whether the call behavior exists fraud; the first training data includes a call behavior trajectory labeled with normal behavior trajectory and all fraud behavior trajectories of fraud types; the second model is used to learn a reward function consistent with fraud based on the fraud behavior trajectory; the automatic acquisition of the first reward function avoids the difficulty of manually designing a complex reward function, improves the adaptability of the model, and reduces the dependence on a large amount of labeled data. Further, based on the first reward function and the first training data, the first model is updated and trained to obtain an updated first model. The introduction of the first reward function of the application enables the first model to more flexibly cope with different call behaviors and analysis requirements, so as to identify the call behavior based on the updated first model and output whether the call behavior exists fraud, thereby significantly improving the accuracy of telecom fraud detection.
[0053] Embodiment two
[0054] The technical scheme of the embodiment verifies the telecom fraud detection method in the foregoing embodiment based on the foregoing embodiment. Table 1 is a data sample statistic of this verification, and Table 2 is a classification result statistic, that is, a verification result. The specific verification process includes the following contents:
[0055] The application adopts a graph neural network model (first model) to identify fraud behavior. Further, in order to verify the accuracy of the first model of the application, four models, including a graph convolution network (GCN) model, a model (GAT) model, a relational graph convolution network (RGCN) model, and a graph isomorphism network (GIN) model, are selected as comparison objects. The first model is compared with the four models having unique structures and functions, that is, the graph convolution network (GCN) model, the model (GAT) model, the relational graph convolution network (RGCN) model, and the graph isomorphism network (GIN) model. The verification result of the model can be mainly evaluated by three key indicators, that is, accuracy, AUC value, and recall rate, to ensure the accuracy and effectiveness of the model.
[0056] In actual application, the GCN model captures the complex relationship between nodes through graph convolution operation; the GAT model introduces an attention mechanism to optimize the weighted aggregation of neighbor nodes, thereby improving the robustness of the model to noise neighbors; the RGCN model is specially designed for processing multi-relation graph data; and the GIN model updates the node representation by aggregating neighbor features, and its performance can be comparable to the WL-test in some cases.
[0057] By considering these evaluation indicators and model characteristics comprehensively, a more comprehensive evaluation and selection of a graph neural network model suitable for a specific task can be made. For example, in the field of fraud detection, a model with high recall rate and high AUC value may be more desirable, as it can more effectively identify fraudulent behavior while maintaining a lower false positive rate.
[0058] As can be seen from Table 2, in this experiment, the first model of the present application exhibits excellent performance, with an accuracy of 0.766, an AUC-PR of 0.778, and a recall rate of 0.733. Compared with other models such as GCN (accuracy 0.560, AUC-PR 0.488, recall rate 0.633), GAT (accuracy 0.658, AUC-PR 0.569, recall rate 0.523), RGCN (accuracy 0.689, AUC-PR 0.650, recall rate 0.618), and GIN (accuracy 0.597, AUC-PR 0.601, recall rate 0.571), there is a significant improvement.
[0059] Table 1: Data sample statistics
[0060] Property Number User number (V) 265411 Total number of relationships (E) 1496654 One side number 365151 Maximum degree of node 4632 Total number of marked fraud 2045 (0.77% of the total)
[0061] Table 2: Classification result statistics
[0062] Model name Accuracy AUC Recall rate GCN 0.560 0.488 0.633 GAT 0.658 0.569 0.523 RGCN 0.689 0.650 0.618 GIN 0.597 0.601 0.571 The model of the application 0.736 0.748 0.703
[0063] By introducing the second model, the present application can dynamically adjust the parameters of the reward function in the learning process, thereby quickly and accurately obtaining an accurate reward function, avoiding the difficulty of manually designing a complex reward function, improving the adaptability of the model, and reducing the dependence on a large amount of labeled data, so that the first model is updated and trained using the second model to more accurately capture complex patterns and associated information in the data. The telecommunication fraud detection method of the present application can more effectively extract features and transmit information when processing complex graph structure data, thereby significantly improving the prediction accuracy and generalization ability of the model. This achievement not only brings a new breakthrough in the field of graph neural networks in theory, but also provides a more powerful tool for solving complex problems in practical application scenarios.
[0064] Example Three
[0065] Figure 2 A structural schematic diagram of a telecommunication fraud detection device provided by the embodiment of the present application is provided. The embodiment can be applicable to the case of telecommunication fraud detection. The telecommunication fraud detection device can be realized in the form of hardware and / or software, and can be configured in any electronic device with network communication function. As shown in the figure, the telecommunication fraud detection device of the present application comprises: Figure 2
[0066] The function determination module 210 is configured to, if it is detected that the accuracy of the first model is less than a preset accuracy, update and train a second model based on first training data to obtain a first reward function; the first model is used to detect whether a call behavior exists fraudulent behavior; the first training data includes a call behavior trajectory labeled with a normal behavior trajectory and all fraudulent behavior trajectories of fraudulent behavior types; the second model is used to learn a reward function consistent with fraudulent behavior based on the fraudulent behavior trajectory;
[0067] The updating module 220 is configured to update and train the first model based on the first reward function and the first training data to obtain an updated first model.
[0068] The detection module 230 is configured to identify the call behavior based on the updated first model, and output whether the call behavior exists fraudulent behavior.
[0069] On the basis of the above embodiment, optionally, the second model includes a target function, the target function is used to describe the probability of the reward function consistent with fraudulent behavior, the first reward function is obtained by updating and training the second model based on the first training data, including:
[0070] Adjusting the target function based on the first training data, when the target value of the target function is greater than a preset value, determining the reward function corresponding to the target function as the first reward function.
[0071] On the basis of the above embodiment, optionally, the target function is expressed by the following formula:
[0072] L(θ)=∑ τ∈D logP(τ∣θ)-logZ(θ);
[0073] Wherein, L(θ) is a target value; D is a fraudulent behavior trajectory in the first training data; P(τ∣θ) is the probability of the fraudulent behavior trajectory τ in the first training data under the parameter θ; Z(θ) is a normalization constant, and the normalization constant Z(θ) is expressed as: Z(θ)=∑ τ∈Ω P(τ∣θ), Ω is a set of all fraudulent behavior trajectories.
[0074] On the basis of the above-mentioned embodiment, optionally, the target function is constrained by a preset constraint condition, the preset constraint condition is used to describe that a first expected value is same as a second expected value; the first expected value is an expected value of a feature function on a fraud behavior trajectory generated by the second model; the second expected value is an expected value of the feature function on the fraud behavior trajectory in the first training data; the feature function is used to describe a state and an action of the fraud behavior; when a target value of the target function is greater than a preset value, the first reward function is constructed by using the feature function corresponding to the target function with the target value greater than the preset value.
[0075] On the basis of the above-mentioned embodiment, optionally, based on the first reward function and the first training data, the first model is updated and trained to obtain an updated first model, comprising:
[0076] Based on a classification loss function and a reward signal of the first reward function, a target loss function of the first model is constructed;
[0077] Based on the first training data and the target loss function, parameters of the first model are adjusted until a loss value of the target loss function is less than a preset loss value, and an updated first model is obtained.
[0078] On the basis of the above-mentioned embodiment, optionally, based on a classification loss function and a reward signal of the first reward function, a target loss function of the first model is constructed, comprising:
[0079] The target loss function is expressed by the following formula:
[0080] L=L1-λf;
[0081] Wherein, L1 is the classification loss function; f is the reward signal of the first reward function; and λ is a regularization coefficient.
[0082] On the basis of the above-mentioned embodiment, optionally, a representation form of the first training data is a graph data structure of a telecommunication network call environment, the graph data structure is composed of nodes and edges, the nodes are two call objects of a call behavior, and the two call objects of the call behavior are connected to become edges; attributes of the nodes at least include basic information and behavior characteristics of the call behavior; attributes of the edges at least include a call frequency, a call duration and a weight of the edge; and the weight of the edge is determined based on the call frequency and the call duration.
[0083] The telecommunication fraud detection device provided in the embodiment of the application can execute the telecommunication fraud detection method provided in any embodiment of the application, and has the corresponding function modules and beneficial effects of the execution method.
[0084] Embodiment four
[0085] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0086] Figure 3 A structural schematic diagram of an electronic device that can be used to implement the fraud detection method of the present application is shown. The electronic device is intended to represent a variety of forms including digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown in the figures, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0087] As shown in Figure 3 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication, wherein the memory stores a computer program executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0088] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0089] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the telecommunication fraud detection method.
[0090] In some embodiments, the telecommunication fraud detection method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the telecommunication fraud detection method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the telecommunication fraud detection method by any other suitable means, such as by means of firmware.
[0091] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0092] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0093] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0095] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0096] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0097] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.
[0098] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method of telecommunications fraud detection, characterized by, The method comprises: If it is detected that the accuracy of the first model is less than a preset accuracy, the second model is updated and trained based on first training data to obtain a first reward function; the first model is used to detect whether a call behavior exists fraud behavior; the first training data comprises a call behavior trajectory labeled with normal behavior trajectory and fraud behavior trajectory of all fraud types; the second model is used to learn a reward function consistent with fraud behavior based on the fraud behavior trajectory; The first model is updated and trained based on the first reward function and the first training data to obtain an updated first model; Based on the updated first model, the call behavior is identified, and whether the call behavior exists fraud behavior is output.
2. The method of claim 1, wherein, The second model comprises a target function used to describe the probability of the reward function consistent with fraud behavior, the second model is updated and trained based on first training data to obtain a first reward function, comprising: The target function is adjusted based on the first training data, and when the target value of the target function is greater than a preset value, the reward function corresponding to the target function is determined as the first reward function.
3. The method of claim 2, wherein, The target function is represented by the following formula: L(θ) = ∑ τ∈D log P(τ | θ) - log Z(θ); wherein L(θ) is a target value; D is a fraud behavior track in the first training data; P(τ|θ) is a probability of a fraud behavior track τ in the first training data under a parameter θ; Z(θ) is a normalization constant, and the normalization constant Z(θ) is expressed as: Z(θ) =∑ τ∈Ω P(τ|θ), and Ω is a set of all fraud behavior tracks.
4. The method of claim 3, wherein, The target function is constrained by a preset constraint condition, and the preset constraint condition is used to describe that a first expected value is the same as a second expected value; the first expected value is an expected value of a feature function on a fraud behavior trajectory generated by the second model; The second expected value is an expected value of a feature function on a fraud behavior trajectory in the first training data; the feature function is used to describe the state and action of fraud behavior; when the target value of the target function is greater than a preset value, the first reward function is constructed by using the feature function corresponding to the target function whose target value is greater than the preset value.
5. The method of claim 1, wherein, The first model is updated and trained based on the first reward function and the first training data to obtain an updated first model, comprising: Based on the classification loss function and the reward signal of the first reward function, a target loss function of the first model is constructed; Based on the first training data and the target loss function, the parameters of the first model are adjusted until the loss value of the target loss function is less than a preset loss value, and an updated first model is obtained.
6. The method of claim 5, wherein, Based on the classification loss function and the reward signal of the first reward function, a target loss function of the first model is constructed, comprising: The target loss function is represented by the following formula: L=L1-λf; Wherein, L1 is the classification loss function; f is the reward signal of the first reward function; λ is a regularization coefficient.
7. The method of claim 1, wherein, The representation form of the first training data is a graph data structure of a telecommunication network call environment, the graph data structure is composed of nodes and edges, the nodes are two call objects of a call behavior, and the two call objects of the call behavior are connected into edges; The attributes of the nodes include at least basic information and behavior characteristics of the call behavior; The attributes of the edges include at least call frequency, call duration and weight of the edges; The weight of the edge is determined based on the call frequency and the call duration.
8. A telecommunications fraud detection apparatus characterized by, The device comprises: The function determining module is configured to, if it is detected that the accuracy of the first model is less than a preset accuracy, update and train a second model based on first training data to obtain a first reward function; the first model is configured to detect whether a call behavior is fraudulent; the first training data includes a call behavior trajectory labeled with a normal behavior trajectory and a fraudulent behavior trajectory of all types of fraudulent behavior; and the second model is configured to learn a reward function consistent with fraudulent behavior based on the fraudulent behavior trajectory. The updating module is configured to update and train the first model based on the first reward function and the first training data to obtain an updated first model. The detecting module is configured to identify a call behavior based on the updated first model and output whether the call behavior is fraudulent.
9. An electronic device, comprising: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the telephonic fraud detection method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the telephonic fraud detection method of any one of claims 1-7 when executed.