Telecommunication fraud identification method and device, equipment, storage medium and product
By constructing a heterogeneous graph of telecommunications users' communication behavior and employing a multi-curvature parameterized exponential mapping technique, node features are mapped to multi-layer hyperbolic graph convolutional layers. Combined with a semi-supervised learning optimization model, the problem of insufficient accuracy and efficiency in existing technologies for telecommunications fraud identification is solved, and efficient identification of telecommunications fraud users is achieved.
Patent Information
- Application Number
- CN202511110265.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-14
AI Technical Summary
Existing graph learning-based telecommunications fraud identification technologies suffer from data feature distortion when processing communication network data that exhibits power-law distribution and hierarchical structure. This leads to insufficient discovery of potential connections between fraudulent users, affecting the accuracy and efficiency of telecommunications fraud identification.
A heterogeneous graph of the first communication behavior of telecommunications users is constructed. The Euclidean features of nodes are mapped to the hyperbolic space of a multi-layer hyperbolic graph convolutional layer through a multi-curvature parameterized exponential mapping. Feature updates are performed within the multi-layer hyperbolic graph convolutional layer. The multi-curvature hyperbolic heterogeneous graph neural network model is optimized by combining semi-supervised learning, and the node fraud probability is output to identify telecommunications fraud users.
It effectively captures the power-law distribution and hierarchical structure characteristics in communication networks, uncovers the complex relationships between nodes, especially the potential connections between fraudulent users, and improves the accuracy and efficiency of fraud identification in telecommunications scenarios.
Smart Images

Figure CN120951093A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, devices, equipment, storage media and products for identifying telecommunications fraud. Background Technology
[0002] The rapid expansion of the telecommunications industry and the dramatic increase in the number of users have led to a surge in telecommunications fraud. These fraudulent activities not only threaten users' privacy and financial security but also significantly impact the operational efficiency and reputation of telecommunications operators.
[0003] Current telecommunications fraud user identification technologies can be broadly categorized as follows: rule-based identification, statistical analysis-based identification, machine learning-based identification, deep learning-based identification, and graph learning-based identification. Among these, graph learning-based identification is particularly suitable for revealing fraud patterns that rely on complex relationships and is well-suited for handling fraud involving multiple users. However, current graph learning-based identification suffers from data feature distortion when processing communication network data that exhibits power-law distributions and hierarchical structures. This leads to insufficient discovery of potential connections between fraudulent users, affecting the accuracy and efficiency of telecommunications fraud identification.
[0004] In summary, improving the accuracy and efficiency of fraud identification in telecommunications scenarios has become a pressing technical problem that needs to be solved in this field. Summary of the Invention
[0005] The main purpose of this application is to provide a method, apparatus, device, storage medium and product for identifying telecommunications fraud, which aims to improve the accuracy and efficiency of fraud identification in telecommunications scenarios.
[0006] To achieve the above objectives, this application proposes a method for identifying telecommunications fraud, which includes: Construct a first communication behavior heterogeneous graph for telecommunications users, which contains nodes of various different types; The first Euclidean features of each node are mapped to the first hyperbolic space of the multi-curvature parameterized hyperbolic graph convolutional layer to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters. Within the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are updated to obtain the embedding matrix; By combining the embedding matrix with a semi-supervised learning approach, the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer are optimized to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model. The algorithm outputs node fraud probabilities based on a trained multi-curvature hyperbolic heterogeneous graph neural network model to identify users who commit telecommunications fraud.
[0007] In one embodiment, the step of updating the features of the first hyperbolic node within the multi-layer hyperbolic graph convolutional layer to obtain the embedding matrix includes: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are sequentially subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing to obtain new first Euclidean features. The new first Euclidean feature is mapped to the first hyperbolic space of the next hyperbolic graph convolutional layer connected to the first hyperbolic graph convolutional layer to obtain the new first hyperbolic node feature. Then, the steps of performing hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing on the first hyperbolic node feature are performed sequentially until the current hyperbolic graph convolutional layer is the last layer, and the output embedding matrix is obtained.
[0008] In one embodiment, the hyperbolic graph convolutional layer includes multiple hyperbolic spaces and multiple Euclidean tangent spaces. The step of sequentially performing hyperbolic transformation, heterogeneous attention neighbor aggregation, and nonlinear activation processing on the first hyperbolic node features in the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer to obtain new first Euclidean features includes: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are transformed by logarithmic mapping to the first Euclidean tangent space and then by exponential mapping to the second hyperbolic space to obtain the second hyperbolic node features. After the second hyperbolic feature is logarithmically mapped to the second Euclidean tangent space and heterogeneous attention neighbor aggregation is performed, it is mapped to the third hyperbolic space to obtain the third hyperbolic node feature; The third hyperbolic feature is nonlinearly activated by mapping it to the third Euclidean tangent space using a logarithmic method, resulting in a new first Euclidean feature.
[0009] In one embodiment, the step of mapping the second hyperbolic feature to a second Euclidean tangent space via logarithmic mapping for heterogeneous attention neighbor aggregation, and then mapping it to a third hyperbolic space to obtain the third hyperbolic node feature includes: The second hyperbolic feature is mapped to the second Euclidean tangent space using a logarithmic method to obtain the second communication behavior heterogeneous graph; The second communication behavior heterogeneous graph is divided into subgraphs according to the triples of source node type, edge type, and target node type; Neighbor aggregation is performed on the nodes in each of the subgraphs, and the subgraphs after the neighbor aggregation operation are merged to obtain a third communication behavior heterogeneous graph. The heterogeneous graph of the third communication behavior is mapped to the third hyperbolic space to obtain the third hyperbolic node features.
[0010] In one embodiment, the step of performing neighbor aggregation operation on nodes in each of the subgraphs includes: For each subgraph, determine the source node and the target node in the subgraph; A dynamic attention mechanism is adopted to calculate the relationship-aware attention weights based on the linear transformation results between the source node and the target node; Based on the attention weights, the target node is guided to perform a neighbor aggregation operation to complete the feature update of the target node and obtain a new subgraph.
[0011] In one embodiment, the step of constructing a first heterogeneous graph of communication behavior for telecommunications users includes: Construct a first communication behavior heterogeneous graph of telecommunications users. The first communication behavior heterogeneous graph contains multiple types of nodes and edges, wherein the nodes represent telecommunications users and the edges represent communication behaviors between telecommunications users. The attribute information of each node is characterized to obtain a feature matrix; The topological connections between the nodes are recorded using an adjacency matrix.
[0012] In one embodiment, the step of optimizing the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layers by combining the embedding matrix with a semi-supervised learning approach to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model includes: The embedding matrix is logarithmically mapped to the target Euclidean tangent space to obtain the target Euclidean feature matrix; The target Euclidean feature matrix is transformed using a multilayer perceptron, and a normalization function is used to generate a prediction distribution matrix of nodes belonging to the fraud category. Based on the labeled fraudulent user tags and the predicted distribution matrix, the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer are iteratively updated using the cross-entropy loss function to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model.
[0013] Furthermore, to achieve the above objectives, this application also proposes a telecommunications fraud identification device, which includes: The heterogeneous graph construction module is used to construct a first communication behavior heterogeneous graph of telecommunications users, which contains nodes of various different types. The feature mapping module is used to map the first Euclidean features of each node to the first hyperbolic space of the multi-curvature parameterized hyperbolic graph convolutional layer to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters. The feature update module is used to update the features of the first hyperbolic node within a multi-layer hyperbolic graph convolutional layer to obtain an embedding matrix; The parameter optimization module is used to optimize the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer by combining the embedding matrix through semi-supervised learning, so as to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model. The fraud detection module is used to output the fraud probability of nodes based on a trained multi-curvature hyperbolic heterogeneous graph neural network model in order to identify users who commit telecommunications fraud.
[0014] In addition, to achieve the above objectives, this application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the telecommunications fraud identification method described above.
[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the telecommunications fraud identification method described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the telecommunications fraud identification method described above.
[0017] This application proposes a method for identifying telecommunications fraud. The method involves constructing a first communication behavior heterogeneous graph of telecommunications users, which contains nodes of various types. The first Euclidean features of each node are mapped to the first hyperbolic space of a multi-layer hyperbolic graph convolutional layer using a multi-curvature parameterized exponent to obtain first hyperbolic node features, where different types of nodes are assigned different curvature parameters. Within the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are updated to obtain an embedding matrix. The model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer are optimized using a semi-supervised learning approach combined with the embedding matrix, resulting in a trained multi-curvature hyperbolic heterogeneous graph neural network model. Based on the trained multi-curvature hyperbolic heterogeneous graph neural network model, the node fraud probability is output to identify telecommunications fraud users.
[0018] In summary, this application constructs a first heterogeneous graph of communication behavior containing multiple types of nodes and employs a multi-curvature parameterized exponential mapping technique to map the Euclidean features of nodes to hyperbolic space. This effectively captures the power-law distribution and hierarchical structure characteristics of communication data in the communication network, avoiding data feature distortion. Simultaneously, the design of multi-layer hyperbolic graph convolutional layers can uncover complex relationships between nodes, particularly the potential connections between fraudulent users. Finally, the multi-curvature hyperbolic heterogeneous graph neural network model, optimized based on semi-supervised learning, can efficiently predict node fraud probabilities and accurately identify telecommunications fraud users. This effectively solves the problem of insufficient accuracy and efficiency of current graph learning-based identification techniques when processing complex communication network data, significantly improving the accuracy and efficiency of fraud identification in telecommunications scenarios. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating an embodiment of the telecommunications fraud identification method of this application. Figure 2 This is a schematic diagram of triplet neighbor aggregation provided in Embodiment 2 of the telecommunications fraud identification method of this application; Figure 3 This is a schematic diagram of the telecommunications fraud identification process provided in Embodiment 2 of the telecommunications fraud identification method of this application; Figure 4 This is a schematic diagram of hyperbolic heterogeneous graph convolution provided in Embodiment 2 of the telecommunications fraud identification method of this application; Figure 5 This is a schematic diagram of the module structure of the telecommunications fraud identification device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the telecommunications fraud identification method in this application embodiment.
[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0025] The main solution of this application embodiment is as follows: A first communication behavior heterogeneous graph of telecommunications users is constructed, containing nodes of various types; the first Euclidean features of each node are mapped to the first hyperbolic space of a multi-layer hyperbolic graph convolutional layer through a multi-curvature parameterized exponent to obtain first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters; the first hyperbolic node features are updated within the multi-layer hyperbolic graph convolutional layer to obtain an embedding matrix; the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing multi-layer hyperbolic graph convolutional layers are optimized using a semi-supervised learning method combined with the embedding matrix to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model; and the node fraud probability is output based on the trained multi-curvature hyperbolic heterogeneous graph neural network model to identify telecommunications fraud users.
[0026] This application provides a solution that constructs a first heterogeneous graph of communication behavior containing multiple types of nodes and uses a multi-curvature parameterized exponential mapping technique to map the Euclidean features of the nodes to hyperbolic space. This effectively captures the power-law distribution and hierarchical structure characteristics of communication data in the communication network, avoiding data feature distortion. Simultaneously, the design of multi-layer hyperbolic graph convolutional layers can uncover complex relationships between nodes, especially potential associations between fraudulent users. Finally, the multi-curvature hyperbolic heterogeneous graph neural network model, optimized based on semi-supervised learning, can efficiently predict the probability of node fraud and accurately identify telecommunications fraud users. This effectively solves the problem of insufficient accuracy and efficiency of current graph learning-based identification techniques when processing complex communication network data, and effectively improves the accuracy and efficiency of fraud identification in telecommunications scenarios.
[0027] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a server, tablet computer, or personal computer, or an electronic device capable of performing the above functions. The following description uses a telecommunications fraud identification terminal as an example to illustrate this embodiment and the subsequent embodiments.
[0028] Based on this, embodiments of this application provide a method for identifying telecommunications fraud, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the telecommunications fraud identification method of this application.
[0029] In this embodiment, the telecommunications fraud identification method includes steps S10 to S50: Step S10: Construct a first communication behavior heterogeneous graph for telecommunications users. The first communication behavior heterogeneous graph contains nodes of various different types. A first communication behavior heterogeneous graph is constructed for telecommunications users. This first communication behavior heterogeneous graph comprehensively integrates various communication behavior data of telecommunications users. The heterogeneous graph contains multiple types of nodes and different types of edges, which can comprehensively and meticulously depict the complexity and diversity of user communication behavior.
[0030] In one feasible embodiment, step S10 may include steps S101 to S103: Step S101: Construct a first communication behavior heterogeneous graph of telecommunications users. The first communication behavior heterogeneous graph contains multiple types of nodes and edges, where nodes represent telecommunications users and edges represent communication behaviors between telecommunications users. It's important to note that a graph is a classic data structure composed of nodes and edges. Compared to other data types, a key characteristic of graph data is its ability to intuitively represent complex entities in the real world and their interrelationships. Graph data is generally categorized into heterogeneous graphs and homogeneous graphs based on the number and types of nodes and edges. As a more common and general form, heterogeneous graphs contain multiple types of nodes and edges, whose features may belong to different semantic spaces. A heterogeneous graph can be represented as... ,in Represents the set of all nodes in the graph. denoted as the number of nodes in the graph. Let be the set of edges in the graph. For type mapping functions, and These represent the sets of node and edge types, respectively. hour, This is called a heterogeneous graph. For example, in a telecommunications operator scenario, nodes may include different types such as landline users, mobile phone users, and enterprise users. Similarly, the edges between nodes may also contain various types, such as phone calls, SMS messages, and business payments. And when... hour, It is a homogeneous graph. Characteristic matrix. The matrix consists of the initial features of the nodes. Representative node The features, where R represents the set of real numbers, and the feature dimension is... With the set of edges Correspondingly, graph structures also typically use adjacency matrices. Used to record the topological connections between nodes in a graph. If a node... and nodes If there is an edge between them, then Conversely .
[0031] A heterogeneous graph structure is defined to contain various types of nodes and edges. Nodes are the basic units in a heterogeneous graph, representing different types of telecommunications users, such as fixed-line users, mobile phone users, and enterprise users. Each node contains the attribute information of the corresponding telecommunications user. Edges represent communication behaviors between telecommunications users, such as phone calls, text messages, and business payments. Edges connect different nodes, forming a complex relationship network between users.
[0032] Step S102: Perform feature processing on the attribute information of each node to obtain the feature matrix; For the attribute information contained in the node corresponding to each telecom user, attribute classification, feature extraction and transformation are adopted to convert the attribute information into numerical feature vectors. The feature vectors of each node are integrated into a feature matrix for subsequent graph convolutional layer processing and model training.
[0033] For example, in a feasible implementation scenario, in order to transform various user attribute information into a numerical feature matrix that the model can process, this embodiment first classifies user attributes and formulates corresponding feature processing strategies for each attribute type. In the actual scenario of telecommunications operators, user attributes can be roughly divided into the following categories: Numerical characteristics: such as number of calls, call duration, number of text messages, data traffic consumption, etc.
[0034] Classification features include user level, user gender, account activity level, account security level, and account authentication status (whether real-name authentication is required).
[0035] Text features include user profiles, customer service chat logs, device malfunction reports, and users' public comments on social media.
[0036] Numerical features can be directly used for model training, but they need to be standardized or normalized to ensure consistent scale across different features. For categorical features, this embodiment uses tools such as scikit-learn (a machine learning library) to convert categorical features into one-hot encoding, making them suitable for machine learning models. For textual attributes, this embodiment uses the BERT model for word embedding, converting textual information into high-dimensional dense vectors. Then, feature vectors of all types are concatenated to form the complete representation of a node, and finally, the feature vectors of all nodes are integrated to generate a complete feature matrix. This feature matrix serves as the input data for a multi-curvature hyperbolic heterogeneous graph neural network model.
[0037] Step S103: Record the topological connection relationship of each node through the adjacency matrix to construct the first heterogeneous graph of communication behavior of telecommunications users.
[0038] The adjacency matrix records the topological connections between nodes to ultimately construct a heterogeneous graph of telecommunications users' communication behavior. The adjacency matrix is a matrix that represents the connection relationships between nodes in the graph. The element values of the matrix indicate whether there are edges between nodes and the weight of the edges. Specifically, the corresponding adjacency matrix can be constructed according to the actual communication behavior between telecommunications users to accurately record the topological connections between nodes.
[0039] By combining the defined heterogeneous graph structure, feature matrix, and adjacency matrix, a heterogeneous graph of telecommunications user communication behavior containing rich structural and attribute information can be finally constructed, namely the first heterogeneous graph of communication behavior.
[0040] Step S20: The first Euclidean features of each node are mapped to the first hyperbolic space of the multi-layer hyperbolic graph convolutional layer through exponential mapping with multi-curvature parameterization to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters. It should be noted that the Euclidean features of a node refer to the node attributes or characteristics represented in Euclidean space. In Euclidean space, features are usually represented as vectors, with each component of the vector representing the node's attributes or measurements in different dimensions. Hyperbolic geometry is a type of non-Euclidean geometry, and conclusions of traditional Euclidean geometry, such as the parallel postulate, do not hold in hyperbolic geometry. Two-dimensional hyperbolic geometry is also known as the hyperbolic plane or saddle plane, and its main characteristic is that the Gaussian curvature at each point on the plane is a negative constant, where curvature describes the degree of bending in the neighborhood. Hyperbolic space refers to the case when hyperbolic geometry is extended to three dimensions or more. By definition, it belongs to a special type of Riemannian fluid, where the sectional curvature at each point in space is always negative. The first hyperbolic space refers to the first hyperbolic space in the convolutional layer of a hyperbolic graph. Hyperbolic node features refer to the node attributes or characteristics represented in hyperbolic space. In hyperbolic space, the distances and relationships between nodes no longer follow the rules of Euclidean geometry, but instead present a more complex and hierarchical structure.
[0041] It should also be noted that in this embodiment, the multi-curvature hyperbolic heterogeneous graph convolutional model consists of multiple hyperbolic graph convolutional layers. Considering the data heterogeneity issue, this embodiment designs and employs a multi-curvature approach in the exponential and logarithmic mapping processes. This involves setting different hyperbolic curvatures for different types of nodes, thereby achieving the effect of mapping different types of nodes to different hyperbolic model spaces in each spatial transformation. These curvatures can be used as model parameters, and the optimal spatial mapping relationship can be obtained through model training and optimization.
[0042] The first Euclidean features of each node are mapped to the first hyperbolic space of the multi-layer hyperbolic graph convolutional layer through a multi-curvature parameterized exponential mapping technique, thereby obtaining the first hyperbolic node features. In this mapping process, different curvature parameters are assigned to different types of nodes to ensure that the hierarchical structure and power-law distribution characteristics between nodes can be captured more accurately.
[0043] In one feasible embodiment, step S20 may include steps S201-S202: Step S201: For each type of node, select the corresponding curvature parameter according to the node type; It should be noted that curvature is a crucial parameter in hyperbolic graph neural network models. Since heterogeneous graphs contain various node types, this embodiment establishes a one-to-one correspondence between node type and curvature to utilize type semantic information. Different curvatures are set for different types of nodes, allowing different curvatures to be used as parameters during each spatial mapping and computation within hyperbolic space. This enables the mapping of different node types (e.g., fixed-line users, mobile users, and enterprise users in a telecom operator scenario) onto different hyperbolic geometric models, thus leveraging the node's inherent type information. By setting curvature as a model parameter, this embodiment can optimize the curvature through model learning to obtain the optimal curvature value.
[0044] For different types of nodes, such as fixed-line users, mobile phone users, and enterprise users, each has its own corresponding node type. The curvature parameter is selected according to the node type. This curvature parameter affects the mapping method and distribution pattern of node features in hyperbolic space, which in turn affects the processing effect of subsequent graph convolutional layers and the model's recognition ability.
[0045] Specifically, different node types require different curvature parameters to accurately represent their position and structure in hyperbolic space. For example, core nodes may be located in the central region of hyperbolic space, while edge nodes may be located in the edge region of space. The choice of curvature parameters can affect the distance and connection between nodes. In this embodiment, initial curvature parameters can be pre-set for each type of node based on experience and experimental data. During model training, the curvature parameters are continuously adjusted according to the calculation results of the loss function.
[0046] Step S202: Perform exponential mapping on each node according to each curvature parameter to map the first Euclidean feature of each node to the first hyperbolic space of the multi-layer hyperbolic graph convolutional layer, and obtain the first hyperbolic node feature.
[0047] By using the selected curvature parameters, the first Euclidean feature of each node is mapped to the first hyperbolic space of the multi-layer hyperbolic graph convolutional layer through exponential mapping. This fully utilizes the characteristics of the hyperbolic space to better capture the hierarchical structure and power-law distribution relationship between nodes, thereby obtaining first hyperbolic node features rich in structural information.
[0048] For example, in the scenario of telecommunications operators, nodes may have different attribute characteristics. For instance, landline phones may be associated with some addresses, while the node attribute characteristics of enterprise users may contain some special enterprise or organization information. For different types of nodes in a heterogeneous graph, this embodiment uses different curvatures corresponding to different types of nodes as parameters to perform spatial transformations on each type of node, completing the initial feature mapping from the Euclidean space to the hyperbolic space of the heterogeneous graph nodes. That is, mapping the first Euclidean feature of each node to the first hyperbolic space of the multi-layer hyperbolic graph convolutional layer. This process can be expressed as formula (1):
[0049] In the above formula, this embodiment uses the feature matrix of the node. (i.e., the first Euclidean feature), the hidden features of the nodes in layer 0 are obtained through exponential mapping. (i.e., the first hyperbolic node feature). For any node type Its corresponding curvature can be expressed as Its negative reciprocal For exponential mapping The curvature used in the reference point The extreme points of the hyperbolic model are represented by . After the initial feature transformation, the node features in the resulting hyperbolic space can be used as the raw input for subsequent stacked multi-layer hyperbolic graph convolutional layers.
[0050] Step S30: In the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are updated to obtain the embedding matrix; Within the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features of each node are iteratively updated and deeply fused to form an overall node feature representation, namely the embedding matrix, which can effectively characterize the complex relationships between nodes.
[0051] Step S40: By combining semi-supervised learning with the embedding matrix, the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers are optimized to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model. A semi-supervised learning approach is adopted, and the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers are optimized by combining the embedding matrix. Specifically, the model's ability to identify fraudulent users is continuously improved by minimizing the loss function until a well-trained multi-curvature hyperbolic heterogeneous graph neural network model is obtained. The optimized model parameters include curvature parameters.
[0052] In one feasible embodiment, step S40 may include steps S401 to S403: Step S401: Map the embedding matrix to the target Euclidean tangent space using a logarithmic method to obtain the target Euclidean feature matrix; Since the embedding matrix is obtained in hyperbolic space, it needs to be mapped back to Euclidean space to facilitate subsequent processing using traditional Euclidean space methods. Thus, a logarithmic mapping technique is used to map the embedding matrix to the target Euclidean tangent space, thereby obtaining the target Euclidean feature matrix. This mapping process preserves the structural information in hyperbolic space, while enabling the feature matrix to be effectively processed and analyzed in Euclidean space.
[0053] It should be noted that Euclidean space refers to a linear space with Euclidean metric, while Euclidean tangent space is a local approximation space describing a point on a manifold (a topological space that is locally similar to Euclidean space). Within a local range, Euclidean tangent space is isomorphic to Euclidean space and is used to describe the local geometric properties of the manifold at that point.
[0054] Step S402: Use a multilayer perceptron to transform the dimension of the target Euclidean feature matrix, and use a normalization function to generate a prediction distribution matrix of nodes belonging to the fraud category. The target Euclidean feature matrix is transformed using a multilayer perceptron (MLP), and the predicted distribution matrix of each node belonging to the fraud category is obtained by using the softmax (normalization) function.
[0055] Step S403: Based on the labeled fraudulent user tags and the prediction distribution matrix, the cross-entropy loss function is used to iteratively update the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers, so as to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model.
[0056] Using known fraudulent user labels as supervisory information, the predicted distribution matrix is compared to calculate the cross-entropy loss function, which measures the degree of difference between the predicted result and the true label. Then, the parameters of the network model containing multiple hyperbolic graph convolutional layers are iteratively updated through backpropagation and gradient descent optimization algorithms to minimize the cross-entropy loss function.
[0057] After multiple iterations, the model parameters gradually converge to the optimal value, thus obtaining a well-trained multi-curvature hyperbolic heterogeneous graph neural network model. This model can be used to predict the fraud probability of nodes, i.e., to identify fraudulent users.
[0058] For example, taking node classification as an example, in heterogeneous graph scenarios, the model typically uses the labels of a certain type of node to further classify and distinguish the node categories. The output embedding matrix of the last layer of the hyperbolic graph convolutional network is... Where L represents the number of layers in the convolutional network, R represents the set of real numbers, and N represents the number of samples. To represent the feature dimension, this embodiment first uses a logarithmic mapping to map it onto the tangent plane, and then uses a layer... Complete dimensional transformation and utilize The function obtains the probability distribution matrix of each node's label. In the formula This represents the number of label categories. This embodiment then uses cross-entropy loss to calculate the prediction distribution matrix. With actual label The differences between them. The loss function of the model is shown in Equation (2):
[0059] By calculating the loss function, this embodiment can evaluate the model performance and update parameters, including type-corresponding curvature and those included in the graph neural network model, thereby optimizing the model. Overall, the model utilizes a small amount of labeled sample category information (fraudulent or legitimate nodes) to measure its accuracy in fraudulent user identification scenarios and updates and optimizes the model parameters. After optimization, the model is applied in a real-world task.
[0060] Step S50: Output node fraud probability based on the trained multi-curvature hyperbolic heterogeneous graph neural network model to identify users involved in telecommunications fraud.
[0061] Based on a well-trained hyperbolic heterogeneous graph neural network model, fraud probability prediction is performed on nodes in telecommunications networks. By setting reasonable probability thresholds according to actual application scenarios, potential fraudulent users can be accurately identified, thus enabling telecommunications fraud prevention and control.
[0062] This embodiment provides a method for identifying telecommunications fraud. By constructing a first heterogeneous graph of communication behavior containing multiple types of nodes, and employing a multi-curvature parameterized exponential mapping technique to map the Euclidean features of the nodes to hyperbolic space, it effectively captures the power-law distribution and hierarchical structure characteristics of communication data in the communication network, avoiding data feature distortion. Simultaneously, the design of multi-layer hyperbolic graph convolutional layers can uncover complex relationships between nodes, especially potential associations between fraudulent users. Finally, the multi-curvature hyperbolic heterogeneous graph neural network model, optimized based on semi-supervised learning, can efficiently predict the probability of node fraud and accurately identify telecommunications fraud users. This effectively solves the problem of insufficient accuracy and efficiency of current graph learning-based identification techniques when processing complex communication network data, and significantly improves the accuracy and efficiency of fraud identification in telecommunications scenarios.
[0063] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, step S30 may include steps S301~S302: Step S301: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are sequentially subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing to obtain new first Euclidean features. In any hyperbolic graph convolutional layer of each hyperbolic graph convolutional layer, the received first hyperbolic node features are sequentially subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation, and nonlinear activation to obtain new first Euclidean features.
[0064] Specifically, taking the first hyperbolic graph convolutional layer as an example, firstly, a hyperbolic transformation is performed on the input first hyperbolic node features. This transformation utilizes the characteristics of hyperbolic space to linearly combine and adjust the node features to better capture the complex relationships between nodes. Then, through the heterogeneous attention neighbor aggregation mechanism, the association information between different types of nodes is weighted and fused, enabling the model to dynamically focus on node features that are more important for fraud detection. Finally, a nonlinear activation function is applied to perform a nonlinear transformation on the aggregated features to enhance the model's expressive power, thereby obtaining a new first Euclidean feature. This first Euclidean feature retains the structural information of the nodes in hyperbolic space and incorporates the association information and nonlinear characteristics between nodes.
[0065] In one feasible embodiment, the hyperbolic graph convolutional layer includes multiple hyperbolic spaces and multiple Euclidean tangent spaces, and step S301 may include steps S3011~S3013: Step S3011: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are transformed by logarithmic mapping to the first Euclidean tangent space and then transformed by exponential mapping to the second hyperbolic space to obtain the second hyperbolic node features. It should be noted that during the encoding process of each hyperbolic graph convolutional layer, the node features first need to be hyperbolic mapped, that is, the hyperbolic space features are mapped back to Euclidean space through logarithmic mapping, and then the linear transformation is performed before returning to the hyperbolic space through exponential mapping.
[0066] First, using logarithmic mapping, the features of the first hyperbolic node are mapped from the current hyperbolic space to the first Euclidean tangent space. This first Euclidean tangent space is the first Euclidean tangent space in the current hyperbolic graph convolutional layer. In the Euclidean tangent space, since the spatial characteristics are simpler and easier to process, these features can be linearly transformed. After the linear transformation, the processed features are mapped back to the second hyperbolic space, which is the second hyperbolic space in the current hyperbolic graph convolutional layer, thus obtaining the second hyperbolic node features.
[0067] For example, in heterogeneous graphs, the hyperbolic transformation process is also performed for different node types, i.e., on the 1st... Hyperbolic graph convolutional layer pairs of type When a node undergoes a hyperbolic transformation, its curvature can be expressed as: The hyperbolic transformation of all nodes of this type can be expressed by the following formula (3), where The calculation is given by formula (4).
[0068]
[0069]
[0070] in, This represents the embedding representation of a node of type n_type after hyperbolic transformation in the l-th hyperbolic graph convolutional layer. It is a hyperbolic transform function that performs a hyperbolic transform on the embedded representation of the input. This represents the embedding representation of a node of type n_type in the (l-1)th hyperbolic graph convolutional layer. Indicates curvature. Exponential mapping and logarithmic mapping The curvature used in k represents the pole of the hyperbolic model, which can also be regarded as the origin of the mapping. W represents the weight matrix, and H represents the embedding matrix of the input. k and W can be adjusted during model training.
[0071] Step S3012: After the second hyperbolic feature is logarithmically mapped to the second Euclidean tangent space for heterogeneous attention neighbor aggregation, it is mapped to the third hyperbolic space to obtain the third hyperbolic node feature; Using logarithmic mapping, the second hyperbolic features are mapped to the second Euclidean tangent space. In the Euclidean tangent space, a heterogeneous attention neighbor aggregation mechanism is applied to weightedly fuse the association information between different types of nodes. After completing the heterogeneous attention neighbor aggregation, the processed features are mapped back to the third hyperbolic space, which is the third hyperbolic space in the current hyperbolic graph convolutional layer, thus obtaining the third hyperbolic node features.
[0072] In one feasible embodiment, step S3012 may include steps a10~a40: Step a10: The second hyperbolic feature is mapped to the second Euclidean tangent space through logarithmic mapping to obtain the second communication behavior heterogeneous graph; Using logarithmic mapping techniques, the second hyperbolic feature is mapped from the current hyperbolic space to the second Euclidean tangent space, which is the second Euclidean tangent space in the current hyperbolic graph convolutional layer. The second communication behavior heterogeneous graph is obtained in the Euclidean tangent space.
[0073] Step a20: Divide the heterogeneous graph of the second communication behavior into subgraphs according to the triples of source node type, edge type, and target node type; Based on different combinations of source node type, edge type, and target node type, the heterogeneous graph of the second communication behavior is divided into multiple subgraphs. Each subgraph represents the association between nodes and edges of a specific type. Then, a neighbor aggregation operation is performed in each subgraph. By aggregating the feature information of neighbor nodes, the feature representation of nodes in the subgraph is updated and enriched, thereby obtaining the specific association patterns between nodes and edges of different types.
[0074] Step a30: Perform neighbor aggregation operation on the nodes in each subgraph, and merge the subgraphs after the neighbor aggregation operation to obtain the third communication behavior heterogeneous graph. Neighbor aggregation is performed on the nodes in each subgraph, and the subgraphs after the neighbor aggregation operation are then merged. The fusion method can be simple splicing, weighted summation, or a more complex graph fusion algorithm to integrate the information from different subgraphs into one graph. Through the fusion process, a complete heterogeneous graph of third communication behavior is obtained.
[0075] It's important to note that after the hyperbolic transformation is performed, a neighbor aggregation operation is needed to update the node embeddings in the heterogeneous graph. Traditional heterogeneous graph representation learning models based on meta-paths require obtaining the connections between nodes under different types of meta-paths during the neighbor aggregation process. Conceptually, a meta-path is a predefined path on a specific dataset, linked by multiple node and edge types; semantically, it represents a composite relationship. For example, in a telecommunications service scenario, the path "mobile phone user—fixed-line phone user—mobile phone user" can represent a shared call relationship between different mobile phones calling the same fixed-line phone user.
[0076] Although metapaths can encompass certain semantic information, their design requires manual definition, necessitates prior knowledge while often lacking theoretical support, and is difficult to transfer across different datasets. To avoid the dependence of heterogeneous graph neural network models on metapaths, this embodiment will use <source node type, relation type, target node type> (…<Source,Relation,Target> The inherent data structure of triples in a heterogeneous graph is used as the basic unit for neighbor aggregation, hierarchically aggregating and updating the features of different types of nodes. Specifically, this embodiment selects the inherent structure of triples in the graph as the basis for neighbor aggregation, extracts subgraphs corresponding to different types of triples, and performs neighbor aggregation sequentially within each subgraph, thereby realizing the neighbor aggregation process for all types of nodes in the entire graph. This process can be carried out by... Figure 2 express. Figure 2 In the model, heterogeneous graphs (i.e., heterogeneous graphs of the second communication behavior) are classified according to different types of triples. (Where A represents the adjacency matrix and H represents the embedding matrix) Based on type triples (represented by nodes of different colors in the diagram), the graph is divided into different subgraphs. For example, telecommunications data can be divided into two subgraphs based on two different type triples: <fixed subscribers, calls, mobile phone subscribers> and <enterprise telephone subscribers, calls, mobile phone subscribers>. Within each subgraph, the representations of the nodes are aggregated and updated. Then, through a subgraph fusion process, the representations of all nodes in the entire graph are updated, resulting in the updated embedding matrix. The heterogeneous graph of the third communication behavior is obtained. .
[0077] In one feasible embodiment, the step a30 of "performing neighbor aggregation operation on nodes in each subgraph" may include steps a301 to a303: Step a301: For each subgraph, determine the source node and target node in the subgraph; For each partitioned subgraph, the source node that initiates communication or behavior and the target node that receives communication or behavior are distinguished. By accurately identifying the source node and the target node, the relationship between nodes can be clearly defined.
[0078] Step a302: A dynamic attention mechanism is adopted to calculate the relationship-aware attention weights through the linear transformation results of the source node and the target node; A dynamic attention mechanism is adopted, which can adaptively adjust the importance of the relationship between different nodes. By performing a linear transformation on the features of the source node and the target node, their feature representations are obtained. Then, based on these feature representations, the attention weight between the source node and the target node is calculated. The attention weight reflects the degree of contribution of the source node to the feature update of the target node. It can dynamically perceive the strength of the relationship between nodes, thereby realizing the differentiated aggregation of information of different neighbor nodes.
[0079] Step a303: Based on the attention weight, guide the target node to perform neighbor aggregation operation to complete the feature update of the target node and obtain a new subgraph.
[0080] The target node is guided to perform neighbor aggregation operation based on attention weights. For example, the features of the target node's neighbor nodes are summed according to the attention weights, and then the summation result is fused with the target node's own features to complete the feature update of the target node, thereby obtaining a more accurate and richer node feature representation.
[0081] By updating the features of all target nodes in each subgraph, a new subgraph is finally obtained.
[0082] Step a40: Map the heterogeneous graph of the third communication behavior to the third hyperbolic space to obtain the features of the third hyperbolic node.
[0083] By using the exponential mapping technique, the heterogeneous graph of the third communication behavior is mapped to the third hyperbolic space to obtain the features of the third hyperbolic nodes.
[0084] Step S3013: The third hyperbolic feature is nonlinearly activated by mapping it to the third Euclidean tangent space through logarithmic transformation to obtain a new first Euclidean feature.
[0085] Using logarithmic mapping, the third hyperbolic feature is mapped to the third Euclidean tangent space. In the Euclidean tangent space, a nonlinear activation function is applied to perform a nonlinear transformation on the feature to enhance the model's expressive power, enabling the model to learn more complex nonlinear relationships. After the nonlinear activation is completed, the resulting feature is the new first Euclidean feature, which can be input into the next hyperbolic graph convolutional layer for graph convolution processing.
[0086] It is worth mentioning that on homogeneous graphs, models such as GAT (Graph Attention Network) calculate the importance coefficients of neighboring nodes to each other through a single-layer feedforward network according to formula (5), and use the Softmax function to normalize and obtain the final attention weights.
[0087]
[0088] in, This represents the attention weight of node j towards node i. This represents the normalization of the attention weights for all neighboring nodes j∈N(i) of node i, where N(i) represents the set of neighboring nodes of node i. This represents a nonlinear activation function. This represents a trainable weight vector. This represents a trainable weight matrix. This represents the feature vector of node i. This represents the feature vector of node j. This indicates that the feature vectors of node i and node j after linear transformation are concatenated.
[0089] Building upon this, since nodes of different types (such as mobile phone users, landline users, etc.) on a heterogeneous graph may reside in different semantic spaces, this embodiment proposes a dynamic attention neighbor aggregation method based on type triples. This method calculates attention weights for heterogeneous neighbors, completing the neighbor aggregation process. Before this process, this embodiment first utilizes a logarithmic mapping, using the curvature corresponding to the type as a parameter, to map the representation of the hyperbolic space to the Euclidean tangent space. That is, the second hyperbolic feature is logarithmically mapped to the second Euclidean tangent space, resulting in a second heterogeneous graph of communication behavior.
[0090] Then, within the tangent space, targeting the triplet Let the hidden layer representations of the two endpoints in the previous layer be respectively. and In the The type node is the target node, which aggregates the types of its neighbors. When representing nodes, this embodiment first calculates the dynamic attention weight coefficients according to the following formula:
[0091]
[0092]
[0093] In Equations (6) and (7), K(s) represents a portion of the dynamic attention weight coefficients of the source node s. It is a linear mapping function used for the hidden layer representation of the source node s. After performing a linear transformation, Q(t) represents a portion of the dynamic attention weight coefficients of the target node t. It is a linear mapping function used for the hidden layer representation of the target node t. Perform a linear transformation; in formula (8), Att(s,r,t) represents the attention weight of the source node s to the target node t under relation r. Let N(t) represent the softmax function, which normalizes the attention scores of all neighboring source nodes s of the target node t, such that the sum of the attention weights of all neighboring source nodes s is 1. Let N(t) represent the set of neighboring nodes of the target node t. This represents the transpose of a parameterized matrix, which is a trainable vector parameterized for the relation type r. Represents a non-linear activation function. Represents the weight matrix. This indicates that K(s) and Q(t) are concatenated to obtain a new vector. This is used to scale the attention for different relation types (s,r,t). It is related to the relation triple (s,r,t) and can adjust the attention weights for different relations. d represents the feature dimension. Used to scale attention scores.
[0094] This embodiment is in use Before normalization, a parameter matrix was introduced. The attention score is further calculated using a nonlinear activation function. Traditional dot-product attention mechanisms, when faced with nonlinear, unrelated node features, are also static attention mechanisms, leading to performance degradation. This embodiment, however, achieves dynamic attention weights by changing the attention calculation method.
[0095] After calculating the attention weights Att(s,r,t) between each node's neighbor pairs, this embodiment can apply these attention weights to the source node. The state information is aggregated and updated with neighbors. This process can be represented by the following formula:
[0096]
[0097] Where V(s) represents the state information vector of the source node s, It is a linear mapping function used for the hidden layer representation of the source node s. Perform a linear transformation. Let represent the hidden layer representation of the target node t in the hyperbolic graph convolutional layer at layer l, which is obtained by weighted aggregation of the neighbor node information of the target node t. Att(s,r,t) represents the attention weight of the source node s to the target node t under relation r, and N(t) represents the set of neighbor nodes of the target node t.
[0098] After the neighbor aggregation process is complete, the model performs hyperbolic activation function calculations based on the node type, using the curvature corresponding to different node types in this layer and the next layer as parameters. The result is then passed to the next hyperbolic convolutional layer, and this process is repeated to complete all hyperbolic graph neural network operations.
[0099] Step S302: Map the new first Euclidean feature to the first hyperbolic space of the next hyperbolic graph convolutional layer connected to the first hyperbolic graph convolutional layer to obtain the new first hyperbolic node feature. Then, perform hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing on the first hyperbolic node feature in sequence until the current hyperbolic graph convolutional layer is the last layer to obtain the output embedding matrix.
[0100] Using the exponential mapping technique, the new first Euclidean features are mapped to the first hyperbolic space of the next hyperbolic graph convolutional layer, thereby obtaining the new first hyperbolic node features. The hyperbolic transformation, heterogeneous attention neighbor aggregation, and nonlinear activation operations are performed iteratively.
[0101] Through layer-by-layer iteration and updates, the model can gradually uncover deeper relationships and structural characteristics between nodes. When processing the last hyperbolic graph convolutional layer, the obtained hyperbolic node features will be combined into an embedding matrix, which contains complex relationship information and structural characteristics between nodes.
[0102] In summary, this embodiment constructs a first heterogeneous graph of telecommunications users' communication behaviors and employs multi-curvature parameterized exponential mapping technology to map node features to the first hyperbolic space of a multi-layer hyperbolic graph convolutional layer. This effectively captures the complex hierarchical structure and power-law distribution characteristics among nodes in the telecommunications network. Within the hyperbolic graph convolutional layer, through hyperbolic transformation, heterogeneous attention neighbor aggregation, and layer-by-layer iterative processing of nonlinear activation, in-depth mining and fusion of node features are achieved, enhancing the model's ability to identify fraudulent users. Notably, the introduced dynamic attention mechanism can adaptively adjust the importance of associations between different nodes, further improving the accuracy and efficiency of neighbor aggregation. Finally, by optimizing the network model parameters through semi-supervised learning, the trained multi-curvature hyperbolic heterogeneous graph neural network model can accurately output the node fraud probability, improving the efficiency and accuracy of telecommunications fraud user identification.
[0103] It is worth mentioning that, in order to effectively utilize user characteristics and complex communication behavior data in telecommunications scenarios to achieve efficient and accurate fraud user identification, this embodiment models user communication behavior data as a heterogeneous graph structure, namely, the first communication behavior heterogeneous graph. This heterogeneous graph not only contains different types of nodes (such as users, devices, etc.), but also includes diverse communication behaviors and relationships, fully reflecting the heterogeneity characteristics in telecommunications networks. By using a graph neural network model for feature extraction and fraud user classification, this embodiment can better handle this heterogeneous data. However, traditional graph neural networks based on Euclidean space often cause high data distortion when processing heterogeneous graphs with hierarchical structures or power-law distribution characteristics, making it difficult to accurately capture the complex semantic relationships between nodes. To solve this problem, this embodiment proposes a telecommunications fraud identification method based on a multi-curvature hyperbolic heterogeneous graph neural network, which can effectively reduce data distortion, fully utilize the characteristics of diverse nodes and edges in the heterogeneous graph, and significantly improve the identification effect of fraud users.
[0104] For example, to aid in understanding the implementation flow of the telecommunications fraud identification method obtained by combining the above embodiments, this embodiment provides a simplified flowchart of the telecommunications fraud identification method, as follows: Figure 3 As shown, specifically, the telecommunications fraud identification process includes several modules such as graph input modeling, initial feature mapping, multi-curvature hyperbolic heterogeneous graph convolution operation, and model optimization and update. In this embodiment, the task of fraud user identification is regarded as a semi-supervised node classification task. A corresponding objective loss function is designed, and gradient descent is used to update the parameters of the model. After multiple iterations, the final optimized model is obtained, that is, the trained multi-curvature hyperbolic heterogeneous graph neural network model is obtained, and the model is used for the actual fraud user identification task.
[0105] The multi-curvature hyperbolic heterogeneous graph neural network model includes multiple hyperbolic graph convolutional layers, with the first layer being the most complex. Taking a hyperbolic graph convolutional layer as an example, the calculation process within the layer is illustrated below. Figure 4 As shown.
[0106] First, this embodiment utilizes exponential mapping to transform the heterogeneous graph of the first communication behavior (i.e. Figure 4 In the input graph data, heterogeneous nodes (such as fixed-line users, mobile phone users, and enterprise users in a telecommunications scenario, hereinafter referred to as...) The first Euclidean feature of (representation) is mapped to the second Euclidean feature. Within the first hyperbolic space of the hyperbolic graph convolutional layer. Different curvature parameters are used as model parameters (i.e., ... Figure 4 In , and In this embodiment, nodes of different types are mapped to hyperbolic models with different curvatures. Then, in the... Inside the hyperbolic graph convolutional layer, this embodiment performs hyperbolic transformations according to formula (4) for different types of nodes. Subsequently, based on formulas (11) and (12), heterogeneous attention neighbor aggregation and nonlinear activation function activation are performed in the tangent space using logarithmic mapping. For the heterogeneous attention neighbor aggregation process, this embodiment proposes a triplet-based subgraph aggregation, using the inherent structure of the heterogeneous graph—the type triplet—as the basic unit to achieve neighbor aggregation, thus eliminating the need for manually pre-defining meta-paths as auxiliary structures. The curvature used in the spatial mapping performed inside the hyperbolic graph convolutional layer is based on the curvature in the graph. It means that among them Represents the layer number, The curvatures represent node types and can be learned during model optimization. At the end of each layer, this embodiment calculates the curvature of the next hyperbolic graph convolutional layer according to formula (12). As parameters, the activated features (i.e., the first Euclidean features) are used with an exponential function. Mapped back to hyperbolic space, this data is passed as input to the next hyperbolic graph convolutional layer for subsequent computation. In the final layer of the hyperbolic graph convolutional layer, the multicurvature hyperbolic heterogeneous graph neural network model obtains different embedding matrices for different types of nodes.
[0107] Formulas (4), (11), and (12) are as follows:
[0108]
[0109]
[0110] Among them, in formula (4) It is a hyperbolic transform function that performs a hyperbolic transform on the embedded representation of the input. Logarithmic mapping The curvature used in The extreme points of the hyperbolic model can also be considered as the origin of the mapping, W represents the weight matrix, and H represents the input embedding matrix; in formula (11) Indicates a node A function that aggregates information about neighboring nodes. Represents a node The set of neighboring nodes, Represents a node The neighboring nodes, Indicated by The hyperbolic logarithmic mapping function as a reference point Represents a node The embedded representation; in formula (12) This represents a nonlinear activation function used for transformations and conversions between hyperbolic spaces with different curvatures or different layers. This represents the transformation from the curvature k(l) of layer l to the curvature k(l+1) of layer l+1. Represents a non-linear activation function. Indicates origin The hyperbolic logarithmic mapping function serves as the reference point.
[0111] Using the final embedding results, this embodiment can define a loss function in conjunction with specific downstream tasks to evaluate the embedding quality and optimize the parameters of the updated model.
[0112] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the telecommunications fraud identification method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0113] This application also provides a telecommunications fraud identification device; please refer to... Figure 5 Telecommunications fraud detection devices include: Heterogeneous graph construction module 10 is used to construct a first communication behavior heterogeneous graph of telecommunications users, which contains nodes of various types. Feature mapping module 20 is used to map the first Euclidean features of each node to the first hyperbolic space of the multi-curvature parameterized hyperbolic graph convolutional layer through exponential mapping to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters; The feature update module 30 is used to update the features of the first hyperbolic node within the multi-layer hyperbolic graph convolutional layer to obtain the embedding matrix; The parameter optimization module 40 is used to optimize the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers by combining the embedding matrix with a semi-supervised learning method, so as to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model. Fraud detection module 50 is used to output node fraud probabilities based on a trained multi-curvature hyperbolic heterogeneous graph neural network model in order to identify users who commit telecommunications fraud.
[0114] Optionally, the feature update module 30 is also used for: Within the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are sequentially subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation, and nonlinear activation processing to obtain new first Euclidean features. The new first Euclidean feature is mapped to the first hyperbolic space of the next hyperbolic graph convolutional layer connected to the first hyperbolic graph convolutional layer to obtain the new first hyperbolic node feature. Then, the first hyperbolic node feature is subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing in sequence until the current hyperbolic graph convolutional layer is the last layer, and the output embedding matrix is obtained.
[0115] Optionally, the hyperbolic graph convolutional layer includes multiple hyperbolic spaces and multiple Euclidean tangent spaces, and the feature update module 30 is also used for: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are transformed by logarithmic mapping to the first Euclidean tangent space and then by exponential mapping to the second hyperbolic space to obtain the second hyperbolic node features. After heterogeneous attention neighbor aggregation by mapping the second hyperbolic features to the second Euclidean tangent space through logarithmic mapping, the features are mapped to the third hyperbolic space to obtain the third hyperbolic node features. The third hyperbolic feature is nonlinearly activated by mapping it to the third Euclidean tangent space using a logarithmic method, resulting in a new first Euclidean feature.
[0116] Optionally, the feature update module 30 is also used for: The second hyperbolic feature is mapped to the second Euclidean tangent space through a logarithmic mapping to obtain the second communication behavior heterogeneous graph; The heterogeneous graph of the second communication behavior is divided into subgraphs according to the triples of source node type, edge type, and target node type. Perform neighbor aggregation on the nodes in each subgraph, and then merge the subgraphs after the neighbor aggregation operation to obtain the third communication behavior heterogeneous graph. Mapping the heterogeneous graph of the third communication behavior to the third hyperbolic space yields the features of the third hyperbolic node.
[0117] Optionally, the feature update module 30 is also used for: For each subgraph, determine the source node and the target node in the subgraph; A dynamic attention mechanism is adopted to calculate the relationship-aware attention weights through the linear transformation results between the source node and the target node; Attention weights guide the target node to perform neighbor aggregation operations in order to update the target node's features and obtain a new subgraph.
[0118] Optionally, the heterogeneous graph construction module 10 is also used for: Construct a first communication behavior heterogeneous graph for telecommunications users. The first communication behavior heterogeneous graph contains multiple types of nodes and edges, where nodes represent telecommunications users and edges represent communication behaviors between telecommunications users. The attribute information of each node is characterized to obtain the feature matrix; The topological connections of each node are recorded by an adjacency matrix to construct a heterogeneous graph of the first communication behavior of telecommunications users.
[0119] Optionally, the parameter optimization module 40 is also used for: The embedding matrix is logarithmically mapped to the target Euclidean tangent space to obtain the target Euclidean feature matrix; The target Euclidean feature matrix is transformed using a multilayer perceptron, and a normalization function is used to generate a prediction distribution matrix of nodes belonging to the fraud category. Based on the labeled fraudulent user tags and the prediction distribution matrix, the cross-entropy loss function is used to iteratively update the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers, resulting in a trained multi-curvature hyperbolic heterogeneous graph neural network model.
[0120] The telecommunications fraud identification device provided in this application, employing the telecommunications fraud identification method described in the above embodiments, can improve the accuracy and efficiency of fraud identification in telecommunications scenarios. Compared with the prior art, the beneficial effects of the telecommunications fraud identification device provided in this application are the same as those of the telecommunications fraud identification method provided in the above embodiments, and other technical features in the telecommunications fraud identification device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0121] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the telecommunications fraud identification method in Embodiment 1 above.
[0122] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, complete devices such as multimedia interactive all-in-one machines and touch screen all-in-one machines. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0123] like Figure 6As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0124] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0125] The electronic device provided in this application embodiment employs the telecommunications fraud identification method described in the above embodiments, enabling efficient extraction and display of knowledge from massive, multi-source data. Compared with the prior art, the beneficial effects of the electronic device provided in this application embodiment are the same as those of the telecommunications fraud identification method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment of the telecommunications fraud identification method, and will not be repeated here.
[0126] It should be understood that the various parts disclosed in the embodiments of this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0127] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0128] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the telecommunications fraud identification method in the above embodiments.
[0129] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0130] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0131] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by an electronic device, the electronic device causes the following actions: Constructing a first heterogeneous graph of communication behavior for telecommunications users, the first heterogeneous graph containing nodes of various types; mapping the first Euclidean features of each node to the first hyperbolic space of a multi-layer hyperbolic graph convolutional layer using a multi-curvature parameterized exponent to obtain first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters; updating the first hyperbolic node features within the multi-layer hyperbolic graph convolutional layer to obtain an embedding matrix; optimizing the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer using a semi-supervised learning approach combined with the embedding matrix to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model; and outputting node fraud probabilities based on the trained multi-curvature hyperbolic heterogeneous graph neural network model to identify telecommunications fraud users.
[0132] Computer program code for performing the operations of the embodiments of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0134] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0135] The readable storage medium provided in this application embodiment is a computer-readable storage medium. This medium stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned telecommunications fraud identification method, enabling efficient extraction and display of knowledge from massive amounts of multi-source data. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as those of the telecommunications fraud identification method provided in the above embodiments, and will not be repeated here.
[0136] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the telecommunications fraud identification method described above.
[0137] The computer program product provided in this application embodiment can extract useful information from data generated by information technology systems. Compared with the prior art, the beneficial effects of the computer program product provided in this application embodiment are the same as those of the telecommunications fraud identification method provided in the above embodiments, and will not be repeated here.
[0138] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for identifying telecommunications fraud, characterized in that, The telecommunications fraud identification method includes: Construct a first communication behavior heterogeneous graph for telecommunications users, which contains nodes of various different types; The first Euclidean features of each node are mapped to the first hyperbolic space of the multi-curvature parameterized hyperbolic graph convolutional layer to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters. Within the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are updated to obtain the embedding matrix; By combining the embedding matrix with a semi-supervised learning approach, the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer are optimized to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model. The algorithm outputs node fraud probabilities based on a trained multi-curvature hyperbolic heterogeneous graph neural network model to identify users who commit telecommunications fraud.
2. The telecommunications fraud identification method as described in claim 1, characterized in that, The step of updating the features of the first hyperbolic node within the multi-layer hyperbolic graph convolutional layer to obtain the embedding matrix includes: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the features of the first hyperbolic node are sequentially subjected to hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing to obtain new first Euclidean features. The new first Euclidean feature is mapped to the first hyperbolic space of the next hyperbolic graph convolutional layer connected to the first hyperbolic graph convolutional layer to obtain the new first hyperbolic node feature. Then, the steps of performing hyperbolic transformation, heterogeneous attention neighbor aggregation and nonlinear activation processing on the first hyperbolic node feature are performed sequentially until the current hyperbolic graph convolutional layer is the last layer, and the output embedding matrix is obtained.
3. The telecommunications fraud identification method as described in claim 2, characterized in that, The hyperbolic graph convolutional layer includes multiple hyperbolic spaces and multiple Euclidean tangent spaces. The step of sequentially performing hyperbolic transformation, heterogeneous attention neighbor aggregation, and nonlinear activation processing on the first hyperbolic node features within the first hyperbolic graph convolutional layer in the multi-layer hyperbolic graph convolutional layer to obtain new first Euclidean features includes: In the first hyperbolic graph convolutional layer of the multi-layer hyperbolic graph convolutional layer, the first hyperbolic node features are transformed by logarithmic mapping to the first Euclidean tangent space and then by exponential mapping to the second hyperbolic space to obtain the second hyperbolic node features. After the second hyperbolic feature is logarithmically mapped to the second Euclidean tangent space and heterogeneous attention neighbor aggregation is performed, it is mapped to the third hyperbolic space to obtain the third hyperbolic node feature; The third hyperbolic feature is nonlinearly activated by mapping it to the third Euclidean tangent space using a logarithmic method, resulting in a new first Euclidean feature.
4. The telecommunications fraud identification method as described in claim 3, characterized in that, The step of mapping the second hyperbolic feature to a second Euclidean tangent space via logarithmic mapping for heterogeneous attention neighbor aggregation, and then mapping it to a third hyperbolic space to obtain the third hyperbolic node feature includes: The second hyperbolic feature is mapped to the second Euclidean tangent space using a logarithmic method to obtain the second communication behavior heterogeneous graph; The second communication behavior heterogeneous graph is divided into subgraphs according to the triples of source node type, edge type, and target node type; Neighbor aggregation is performed on the nodes in each of the subgraphs, and the subgraphs after the neighbor aggregation operation are merged to obtain a third communication behavior heterogeneous graph. The heterogeneous graph of the third communication behavior is mapped to the third hyperbolic space to obtain the third hyperbolic node features.
5. The telecommunications fraud identification method as described in claim 4, characterized in that, The step of performing neighbor aggregation operation on nodes in each of the subgraphs includes: For each subgraph, determine the source node and the target node in the subgraph; A dynamic attention mechanism is adopted to calculate the relationship-aware attention weights based on the linear transformation results between the source node and the target node; The attention weights guide the target node to perform neighbor aggregation operations in order to complete the feature update of the target node.
6. The telecommunications fraud identification method as described in claim 1, characterized in that, The step of constructing the first heterogeneous graph of telecommunications user communication behavior includes: Construct a first communication behavior heterogeneous graph of telecommunications users. The first communication behavior heterogeneous graph contains multiple types of nodes and edges, wherein the nodes represent telecommunications users and the edges represent communication behaviors between telecommunications users. The attribute information of each node is characterized to obtain a feature matrix; The topological connections between the nodes are recorded using an adjacency matrix.
7. The telecommunications fraud identification method as described in claim 1, characterized in that, The step of optimizing the model parameters of a multi-curvature hyperbolic heterogeneous graph neural network model containing multiple hyperbolic graph convolutional layers using a semi-supervised learning approach combined with the embedding matrix to obtain a trained multi-curvature hyperbolic heterogeneous graph neural network model includes: The embedding matrix is logarithmically mapped to the target Euclidean tangent space to obtain the target Euclidean feature matrix; The target Euclidean feature matrix is transformed using a multilayer perceptron, and a normalization function is used to generate a prediction distribution matrix of nodes belonging to the fraud category. Based on the labeled fraudulent user tags and the predicted distribution matrix, the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer are iteratively updated using the cross-entropy loss function to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model.
8. A telecommunications fraud identification device, characterized in that, The telecommunications fraud detection device includes: The heterogeneous graph construction module is used to construct a first communication behavior heterogeneous graph of telecommunications users, which contains nodes of various different types. The feature mapping module is used to map the first Euclidean features of each node to the first hyperbolic space of the multi-curvature parameterized hyperbolic graph convolutional layer to obtain the first hyperbolic node features, wherein different types of nodes are assigned different curvature parameters. The feature update module is used to update the features of the first hyperbolic node within a multi-layer hyperbolic graph convolutional layer to obtain an embedding matrix; The parameter optimization module is used to optimize the model parameters of the multi-curvature hyperbolic heterogeneous graph neural network model containing the multi-layer hyperbolic graph convolutional layer by combining the embedding matrix through semi-supervised learning, so as to obtain the trained multi-curvature hyperbolic heterogeneous graph neural network model. The fraud detection module is used to output the fraud probability of nodes based on a trained multi-curvature hyperbolic heterogeneous graph neural network model in order to identify users who commit telecommunications fraud.
9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the telecommunications fraud identification method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the telecommunications fraud identification method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the telecommunications fraud identification method as described in any one of claims 1 to 7.
Citation Information
Cited By
Real-time anomaly detection method for fraud-related webpage based on deep reinforcement learning
CN121193543A
Fraud-related webpage real-time anomaly detection method based on deep reinforcement learning
CN121193543B