A method for detecting financial transaction telecommunication fraud
Patent Information
- Application Number
- CN202610879544.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]然而,上述流程存在根本性局限:它严重依赖历史诈骗案例中总结出的固定特征与模式,本质上是一种“向后看”的匹配机制
本发明实施例提供的一种金融交易电信诈骗检测方法,通过时空图游走,从已知异常用户出发,沿着时间维度追溯行为序列,沿着空间维度遍历社交关系边,即使诈骗者使用不同平台账号,只要在社交网络中存在关联路径(如设备指纹相似、IP关联、行为时间同步),图游走就能将这些账号聚合到同一子图中,突破了单一平台数据孤岛限制,实现跨平台痕迹关联。随后通过动态时间特征捕捉行为动力学特征,对内容变异具有鲁棒性,实现了从关键词匹配到行为序列建模的范式转换,通过拓扑特征
提取用户在网络中的交互行为。再通过多时间戳融合,将离散快照连接为连续演化轨迹,
表征的是用户行为的短期趋势(如:最近3天内,从偶尔发帖变为密集私信,同时拓扑中心性骤升),能够在诈骗链条的早期阶段(如“星星”刚发帖引流时)就检测到异常演化趋势,实现了从事后交易拦截到事前行为预警的转变,不是简单地将图结构文本输入大语言模型LLM,而是用短期演化特征
作为注意力引导,
告诉模型应该关注哪些维度(如:时间演化异常→关注行为顺序语义;拓扑异常→关注关系结构语义),下游编码器对提示词进行再学习和自适应转换,使LLM的通用语义空间对齐到诈骗检测任务空间。当面对新型诈骗时,LLM无需历史样本即可通过语义理解识别风险模式,增强泛化能力,将图结构的拓扑语义与时间演化的动态语义统一编码到同一语义空间,实现跨模态融合,LLM通过语义关联仍能识别新的诈骗模式,实现自适应演化。
Smart Images

Figure CN122736616A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cybersecurity anomaly detection technology, and in particular to a method for detecting telecommunications fraud in financial transactions. Background Technology
[0002] Social networks have become a significant breeding ground for new types of telecommunications fraud. Fraud gangs operate fake accounts across platforms, forge identities, publish misleading content, and engage in targeted interactions, forming a concealed and efficient chain of lead generation and fraud. The multi-dimensional traces left by these actions on social networks, such as account behavior sequences, interactive scripts, content characteristics, and cross-platform connection paths, provide crucial information for real-time identification and interception of fraud. Therefore, building a detection system capable of analyzing social network traces in real time and providing timely warnings of fraud anomalies is of urgent and important practical significance for cutting off fraud chains and protecting users' assets. For example, in April 2025, someone in Nanjing used the social media account "Xingxing" to post recommendations for "reliable purchasing agents," while another WeChat account acted as the purchasing agent; the same person controlled both accounts to commit fraud. In March 2026, a fraud gang in Jiangsu edited and published fake rights protection cases and forged lawyer certificates on platforms such as Douyin and Kuaishou, pushing videos promising "full refund recovery, no success no fee" to attract victims to send private messages. In March 2026, a gang in Shanghai impersonated "successful people" on dating and social media platforms, establishing romantic relationships through long-term online chats before luring victims into fake investment platforms. These examples illustrate that learning about new types of telecom fraud through social networks can be helpful in detecting subsequent unusual transactions.
[0003] The general steps of methods for detecting new types of telecom fraud in financial transactions typically follow this process: First, multi-source data collection and isolated feature extraction are performed. The system independently collects user behavior logs and transaction records from social media platforms and financial transaction systems. On the social media platform side, based on keyword libraries and predefined rules, sensitive words related to fraud, high-frequency contact methods, or specific dialogue patterns are extracted from post content and private message conversations. On the transaction side, structured features such as transaction amount, frequency, time period, and the historical risk level of the recipient are extracted. These features are essentially static and isolated, relying on prior knowledge and usually limited to a single platform or data type. Next, the model training and pattern matching stage based on historical samples is entered. The extracted social features and transaction features are simply concatenated or rule-based, and the supervised learning classification model is trained using previously confirmed fraud case data as positive samples. The core objective of this model is to learn historically defined fraud patterns and feature combinations. Finally, real-time transaction scanning and rule-based threshold warnings are implemented. In application, the system monitors real-time financial transactions, inputting static features extracted from the recent behavior of associated social media accounts along with transaction features into a pre-trained, static model. The model outputs a risk score; if the score exceeds a preset threshold or matches a certain risk rule, an alarm is triggered, attempting to intercept the transaction at the point of sale.
[0004] However, the above process has a fundamental limitation: it relies heavily on fixed characteristics and patterns summarized from historical fraud cases, essentially a "look-back" matching mechanism. Faced with the ever-emerging and rapidly changing fraud methods such as cross-platform collaboration (e.g., posting on social media to drive traffic to WeChat for transactions), fabricating new personas (e.g., impersonating lawyers or successful people), and using evolving rhetoric (e.g., "full refund recovery"), this method is severely inadequate in identifying unknown new fraud patterns, has weak generalization ability, and results in a high rate of false negatives and false negatives. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method for detecting telecommunications fraud in financial transactions, which can identify unknown new fraud patterns, enhance generalization, and reduce losses caused by abnormal information.
[0006] This invention provides a method for detecting telecommunications fraud in financial transactions, comprising the following steps: Enter the user status of the target area where anomaly detection is required; Mapping user state to complex networks Among them For the user set in the target region, Let the set of edges represent the fact that two users know each other in a social network. Starting from the abnormal user, we walk through the spatiotemporal graph to obtain information about other users related to the abnormal user. The system analyzes and updates the order of other users' actions to form dynamic time characteristics. ; Constructing each other user in a complex network The interaction behaviors between users constitute topological features ; Dynamic time characteristics of the same user under a single timestamp and topological features Aggregate the data to obtain its features. ; Data characteristics under all timestamps By performing fusion and evolution, the short-term evolutionary characteristics of users can be obtained. ; Use prompt words to understand complex networks Represented as a large language model with natural language input, the downstream encoder relearns and adaptively transforms the prompt words to obtain the user's encoded features, using short-term evolutionary features. To guide this process, relevance matching and weighted extraction are performed on the encoded features to output a feature representation that incorporates semantic information. ; Will A nonlinear transformation is performed, and the transformation result is mapped to a probability distribution through a normalized exponential transformation to obtain the abnormal probability of all users.
[0007] Optionally, user status includes interactions between social network users.
[0008] Optionally, a spacetime graph walk includes the following steps: Add abnormal users to a queue; Remove the abnormal user and other users associated with them from the queue, and access their connected neighbors; Adjacent users are placed into the visited set and queue. Once a user has been visited, no further sampling will be performed. The sampled data includes the user, the time corresponding to the user, the feature vector corresponding to the user, the user's fuzzy anomaly label, which is either manually marked by the user or initially marked through rule matching, and the user's connection relationship. Output the sequence of users associated with the abnormal user.
[0009] The system analyzes and updates the order of other users' actions to form dynamic time characteristics. This includes the following steps: Generate a memory vector for each other user, with the initial memory vector being a zero matrix; When a new abnormal user and other related users need to be entered, the memory vectors of the other users are updated. The memory vectors contain the user's operation sequence and the corresponding time. The user's action sequence includes the order in which the user likes, comments, and posts. time The feature transformation aligns the dimensions of the memory vector using a single learnable nonlinear mapping function, and the new feature matrix is obtained by concatenating the previous memory vector. The same operation by the same user will be mapped to the same feature matrix through an index, and the feature matrices of all other users will form dynamic temporal features. .
[0010] Optional, constituting topological features This includes the following steps: For another user at present A limited number of samples are taken, selecting users who are friends with the user or have interacted with the user as their neighbor nodes. Corresponding edge This sampling only applies to times less than the current time. Sample users and merge corresponding user features Forming topological features .
[0011] Optionally, obtain data features. By analyzing the dynamic time characteristics of the same user and topological features The average value is obtained by taking the mean.
[0012] Optionally, obtain the user's short-term evolutionary characteristics. This includes the following steps: Data characteristics under all timestamps Perform time series forecasting to obtain new data features at user nodes. While inputting into the next computation stage, the dynamic time features are updated. ; Obtain the data characteristics of each user at the current timestamp. and historical dynamic time characteristics And generate a time location encoding vector; and set the current timestamp The feature matrix is obtained by concatenating it with the time location encoding vector. , will the current timestamp Historical dynamic time characteristics The feature matrix is obtained by concatenating it with the time location encoding vector. and According to the feature matrix , and Obtaining user attention characteristics ; Will The data is input into the historical memory queue, which outputs the features of the historical data. The current data features are then concatenated with the features of the historical memory to obtain the user's short-term evolutionary features. .
[0013] Optionally, the prompts may include: a public description of the target area being collected and the method for constructing user characteristics, user operation information, and timestamps of user state changes; By short-term evolutionary features As Matrix, encoding features as and The matrix input cross-attention mechanism aligns and fuses natural language semantic tasks with graph anomaly detection tasks, outputting a feature representation with fused semantic information. .
[0014] Optionally, it can also output multiple users based on the probability of anomalies from highest to lowest, and output the number of users as required.
[0015] Optionally, this also includes using the output feature representation of the large language model during the training phase. Construct a loss function and optimize the model.
[0016] The technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art: This invention provides a method for detecting telecommunications fraud in financial transactions. Through spatiotemporal graph walking, starting from known abnormal users, it traces behavioral sequences along the time dimension and traverses social relationship edges along the spatial dimension. Even if fraudsters use accounts on different platforms, as long as there are related paths in social networks (such as similar device fingerprints, IP association, or synchronized behavior times), the graph walking can aggregate these accounts into the same subgraph, breaking through the limitations of single-platform data silos and achieving cross-platform trace association. Subsequently, dynamic time features are used... It captures behavioral dynamics features, is robust to content variations, and achieves a paradigm shift from keyword matching to behavioral sequence modeling through topological features. Extract user interactions within the network. Then, through multi-timestamp fusion, connect discrete snapshots into a continuous evolutionary trajectory. It represents short-term trends in user behavior (e.g., a shift from occasional posting to frequent private messages within the last 3 days, accompanied by a sudden increase in topological centrality). It can detect abnormal evolutionary trends in the early stages of a fraud chain (e.g., when "Xingxing" first started posting to attract traffic), achieving a shift from post-transaction interception to pre-event behavioral warning. It doesn't simply input graph-structured text into a large language model (LLM), but rather uses short-term evolutionary features. As a guide to attention The model is told which dimensions to focus on (e.g., temporal evolution anomalies → focus on behavioral sequence semantics; topological anomalies → focus on relational structure semantics). The downstream encoder relearns and adaptively transforms the prompts, aligning the general semantic space of the LLM with the fraud detection task space. When faced with new types of fraud, the LLM can identify risk patterns through semantic understanding without historical samples, enhancing its generalization ability. It unifies the encoding of the topological semantics of the graph structure and the dynamic semantics of temporal evolution into the same semantic space, achieving cross-modal fusion. The LLM can still identify new fraud patterns through semantic association, achieving adaptive evolution. Attached Figure Description
[0017] To more clearly illustrate the embodiments and design schemes of the present invention, the accompanying drawings required for this embodiment will be briefly described below. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a method for detecting telecommunications fraud in financial transactions, provided as part of an embodiment of the present invention; Figure 2 A complete diagram illustrating the memory calculation of a financial transaction telecommunications fraud detection method provided in this embodiment of the invention; Figure 3 A complete illustration of LLM encoding for a method for detecting telecommunications fraud in financial transactions provided by an embodiment of the present invention; Figure 4 A diagram illustrating the loss function and anomaly probability output of a financial transaction telecommunications fraud detection method provided in this embodiment of the invention; Figure 5 This is a comparison chart showing the reduction in data volume and the original data volume after performing a spacetime graph walk method on four real networks—Wikipedia, MOOC, Bitcoin-Alpha, and Bitcoin-OTC—as provided in an embodiment of the present invention. Figure 6This invention provides an accuracy comparison with other methods using only limited data on four real networks: Wikipedia, MOOC, Bitcoin-Alpha, and Bitcoin-OTC. Figure (a) shows the accuracy comparison with other methods using only limited data on the Wikipedia real network; Figure (b) shows the accuracy comparison with other methods using only limited data on the MOOC real network; Figure (c) shows the accuracy comparison with other methods using only limited data on the Bitcoin-Alpha real network; and Figure (d) shows the accuracy comparison with other methods using only limited data on the Bitcoin-OTC real network.
[0019] Figure 7 The detection performance of the model trained on social networks Wikipedia and MOOC provided in this embodiment of the invention on the transaction fraud networks Bitcoin-Alpha and Bitcoin-OTC. Figure 8 This is a flowchart of a method for detecting telecommunications fraud in financial transactions, provided as another embodiment of the present invention. Detailed Implementation
[0020] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.
[0021] The core issue lies in its insufficient generalization ability, making it unable to effectively address unknown new fraud patterns. Due to the scarcity of labeled samples and the continuous evolution of fraud methods, models trained based on fixed features and historical samples struggle to identify unseen fraudulent tactics. Therefore, current detection systems are prone to underreporting or lagging when faced with novel fraud activities involving extremely small sample sizes and continuous evolution. This is precisely the core problem that this invention aims to solve.
[0022] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] First, embodiments of the present invention provide a method for detecting telecommunications fraud in financial transactions, specifically as follows: Figure 1 and Figure 7 As shown, it includes the following steps: Input the user status of the target area to be detected for anomalies, where the target area refers to a real social network platform or a locally built virtual social network environment.
[0024] Mapping user state to complex networks Among them For the user set in the target region, Let the set of edges represent the fact that two users know each other on a social network (whether the users follow each other or have interacted with each other, such as commenting or forwarding each other's posts). If there are already labeled anomalous users, they are retained. Other data is labeled as normal users. Spatiotemporal graph walks are performed starting from the discovered anomalous users or newly emerging anomalous users (users whose accounts have been publicly banned on social media platforms or labeled by researchers). Centered on the node where the anomalous user is located, graph walks are performed on the data of nodes that are directly or indirectly connected to the anomalous node. A finite number of samples are taken. If there are no labeled anomalous nodes in the initial data, the walks are performed on nodes whose degree is less than the average degree of all nodes at the current time to obtain information of other users related to the anomalous user (username, address, signature, posted content, and other public information). Additional data of other users related to the anomalous user is constructed through the walks. This method can retain all anomalous data in the original data, thereby solving the problem of scarce labeled samples.
[0025] In social networks, users' actions such as liking, commenting, and posting are sequential in time, forming a dynamic temporal feature. .
[0026] Constructing each other user in a complex network The interaction behaviors between users constitute topological features Whether it involves a friend relationship or user interaction, such as mutual following, mutual commenting, or mutual sharing. The dimension is the number of user attributes. For example, suppose there is a dataset with 1000 nodes, then... If the value is 1000, and there are 198 user attributes in the dataset, then... The shape of the matrix is (1000, 198). The dimension is 198, and the user's attributes include basic information such as gender and age;
[0027] At a timestamp In the input, a user may have multiple states or action changes, resulting in the final dynamic features of the user. and topological features In cases where there are multiple instances, this invention provides dynamic time features for the same user under a single timestamp. and topological features By aggregating the data, a unique data feature is obtained. ; Data characteristics under all timestamps By performing fusion and evolution, the short-term evolutionary characteristics of users can be obtained. ; Use prompt words to understand complex networks Represented as a large language model with natural language input, the downstream encoder relearns and adaptively transforms the prompt words to obtain the user's encoded features, using short-term evolutionary features. To guide this process, relevance matching and weighted extraction are performed on the encoded features to output a feature representation that incorporates semantic information. ; Will A nonlinear transformation is performed, and the transformation result is mapped to a probability distribution through a normalized exponential transformation to obtain the abnormal probability of all users.
[0028] This invention provides a method for detecting telecommunications fraud in financial transactions. Through spatiotemporal graph walking, starting from known abnormal users, it traces behavioral sequences along the time dimension and traverses social relationship edges along the spatial dimension. Even if fraudsters use accounts on different platforms, as long as there are related paths in social networks (such as similar device fingerprints, IP association, or synchronized behavior times), the graph walking can aggregate these accounts into the same subgraph, breaking through the limitations of single-platform data silos and achieving cross-platform trace association. Subsequently, dynamic time features are used... It captures behavioral dynamics features, is robust to content variations, and achieves a paradigm shift from keyword matching to behavioral sequence modeling through topological features. Extract user interactions within the network. Then, through multi-timestamp fusion, connect discrete snapshots into a continuous evolutionary trajectory. It represents short-term trends in user behavior (e.g., a shift from occasional posting to frequent private messages within the last 3 days, accompanied by a sudden increase in topological centrality). It can detect abnormal evolutionary trends in the early stages of a fraud chain (e.g., when "Xingxing" first started posting to attract traffic), achieving a shift from post-transaction interception to pre-event behavioral warning. It doesn't simply input graph-structured text into a large language model (LLM), but rather uses short-term evolutionary features. As a guide to attention The model is told which dimensions to focus on (e.g., temporal evolution anomalies → focus on behavioral sequence semantics; topological anomalies → focus on relational structure semantics). The downstream encoder relearns and adaptively transforms the prompts, aligning the general semantic space of the LLM with the fraud detection task space. When faced with new types of fraud, the LLM can identify risk patterns through semantic understanding without historical samples, enhancing its generalization ability. It unifies the encoding of the topological semantics of the graph structure and the dynamic semantics of temporal evolution into the same semantic space, achieving cross-modal fusion. The LLM can still identify new fraud patterns through semantic association, achieving adaptive evolution.
[0029] Optionally, user status includes interactions between social network users, such as comments, reposts, following, and whether they are friends. The data can be obtained by manually entering existing private information or by obtaining public information through web crawlers.
[0030] Optionally, a spacetime graph walk includes the following steps: Add abnormal users to a queue; Remove the abnormal user and other users associated with them from the queue, and access their connected neighbors; Adjacent users are added to the visited set and queue; visited users will not be sampled again. The sampled data includes user... The time corresponding to the user The feature vector corresponding to the user User-related tags The user's edge, i.e., the connection relationship. ; Output the user sequence associated with the abnormal user. .
[0031] The specific, complete steps are as follows: Collected graph network data The edge set includes time information, that is, the actual time of the collected edge information. Assuming that when performing a graph walk on a node (user), information of historical nodes and some future node information have been collected. The reason for the existence of some future node information is that graph walks require processing time, which will have an interval with the actual time of edge collection. Therefore, when performing a graph walk on a node, there will be nodes with historical time and some future nodes with time greater than the current node's time. Perform a spatiotemporal graph walk on all nodes in sequence. Each node visited is designated as the current node. Starting from the current node, perform graph walks on nodes in both historical and future time periods using a breadth-first traversal. Add the current node and its connected nodes to the initial set of nodes to be walked. Each time from Take a node and remove it, then add the neighboring nodes that are connected to the taken node to the new node. The system uses a queue and records the corresponding edge information, continuously traversing until it reaches the beginning or end of the graph data, or reaches the user-defined maximum number of traversals. Each time, it removes the original data. Each node is taken out once, which is considered as one walk. The maximum number of walks in this invention is 150.
[0032] The processing flow after inputting the data obtained from the spatiotemporal graph walkthrough does not require processing all the data.
[0033] In this embodiment, reference Figure 2 The operation sequence of other users is statistically analyzed and updated to form dynamic time characteristics. This includes the following steps: Generate a memory vector for each other user, with the initial memory vector being a zero matrix; When a new abnormal user and other related users need to be entered, the memory vectors of the other users are updated. The memory vectors contain the user's operation sequence and the corresponding time. The user's action sequence includes the order in which the user likes, comments, and posts. time The feature transformation aligns the dimensions of the memory vector using a single learnable nonlinear mapping function, GRU algorithm, and then concatenates the previous memory vectors to obtain a new feature matrix. The same operation by the same user will be mapped to the same feature matrix through an index, and the feature matrices of all other users will form dynamic temporal features. Dynamic time characteristics The GRU algorithm is used for dynamic updates over time.
[0034] Optional, constituting topological features This includes the following steps: For another user at present A limited number of samples were taken from users who were friends with each other or had interactive behaviors such as commenting, forwarding, or liking each other's posts. as its neighboring node , corresponding to the edge This sampling only applies to times less than the current time. Sample users and merge corresponding user features Forming topological features .
[0035] Optionally, obtain data features. By analyzing the dynamic time characteristics of the same user and topological features The average value is obtained by taking the mean. For the same node vector, the average value is taken.
[0036] , in, Corresponding user The dynamic time characteristics, Corresponding user Topological features, Corresponding to the number of users sampled, .
[0037] Optionally, obtain the user's short-term evolutionary characteristics. This includes the following steps: refer to Figure 3 Using GRU to analyze data features across all timestamps Perform time series forecasting to obtain new data features at user nodes. While inputting into the next computation stage, the dynamic time features are updated. GRU is widely used in time series forecasting. It dynamically controls the information flow through update and reset gates, selectively retaining or forgetting historical information, thus effectively mitigating this problem and improving the ability to model long-term dependencies, thereby obtaining long-term evolutionary models of data. Specifically, the memory is initialized as an all-zero matrix, where each row corresponds to the feature evolution of a user (node) over time. After node feature aggregation, the obtained node features are input into the GRU algorithm, combined with the node's historical features, to calculate and output node features containing information at the current moment. Multiple features that the same node may generate in a short period are aggregated. By averaging multiple features of the same node, a single representative feature of the node is obtained, which is then updated in the memory as a state record reflecting the user's long-term evolutionary characteristics.
[0038] Obtain the data characteristics of each user at the current timestamp. and historical dynamic time characteristics And generate a time location encoding vector; and set the current timestamp The feature matrix is obtained by concatenating it with the time location encoding vector. , will the current timestamp Historical dynamic time characteristics The feature matrix is obtained by concatenating it with the time location encoding vector. and According to the feature matrix , and Obtaining user attention characteristics In this embodiment, a multi-head temporal attention mechanism is used to correct arbitrary nodes at different timestamps. Attention characteristics (corresponding to the probability that user motivation or behavior is abnormal), nodes exist The attention features at any given moment mainly involve three input feature matrices: Q (the probability that the user himself is an anomalous), K, and V (the probability that other users related to the user are anomalous).
[0039] , in, Indicates the time of input data. Indicates the dimension of the output data. Indicates the outer product. Indicates matrix concatenation. (Multi-Layer Perceptron) represents the multilayer perceptron algorithm.
[0040] , Finally, attention calculation is performed.
[0041] , in, Represents the feature dimension of the K matrix. and Representing nodes respectively The eigenvectors of the corresponding Q and K matrices, This indicates the user's attention characteristics. Indicates the first One user, This indicates the number of users.
[0042] Will The data is input into the historical memory queue, which outputs the features of the historical data. The current data features are then concatenated with the features of the historical memory to obtain the user's short-term evolutionary features. .
[0043] , in, Indicates matrix concatenation. This represents the first attention feature corresponding to the first sample at time t. , This represents the attention feature of the second sample at time t. This represents the attention feature of the third sample at time t. This represents the user's short-term evolutionary characteristics; if none are specified, it is a matrix of all zeros. Attention features for the current data , This represents the multilayer perceptron algorithm.
[0044] This invention captures long-term and short-term anomaly patterns through GRU and memory mechanisms, combined with a spatiotemporal graph walking strategy centered on known anomalies. This enables the anomaly range to be locked in a local area or a few key entities in the early stages, achieving early detection and early localization, thereby improving detection accuracy and reducing the impact of anomaly propagation.
[0045] Optional prompts include: a public description of the target area being collected and methods for constructing user characteristics, user action information, and timestamps of user state changes. See again for details. Figure 3First, the prompt words are input into the large language model and converted into data features. A portion of the prompt words is generated from a complex network. The prompt words include: Data information (corresponding to publicly available introductions of the collected social media platforms and methods for constructing user features); Input information (corresponding to some basic user information); Timestamp (corresponding to timestamps of user state changes); and Nodes (representing users, each user represented by a separate positive integer according to the order of collection). The user's basic information is also the user's actions, such as posting content, commenting / liking / sharing other content, and other publicly queryable actions the user can perform on social media platforms. Simultaneously, the prompt words in the large language model are relearned and adaptively transformed through a downstream encoder to obtain the user's LLM encoded features. Subsequently, the short-term evolution features of the graph structure are used as the Q-matrix, and the LLM encoded features are used as... and A matrix input cross-attention mechanism is used to align and fuse natural language semantic tasks with graph anomaly detection tasks, outputting a feature representation with fused semantic information. For example, some of the prompts in the Wikipedia dataset are:
[0046] Data information: Wikipedia is public dataset is one month of editsmade by edits on Wikipedia pages. It selected the 1,000. most edited pages asitems and editors who made at least 5 edits as users (a total of 8,227users). This generates 157,474 interactions. It converts the edit text into aLIWC-feature vector. Input information: Timestamp 37868.0, Source Node 3 is connected to Node 7049.
[0047] Simultaneously, the LLM's prompts are relearned by the encoder and transformed into a converter specific to the current task. Leveraging the general semantic understanding capabilities of LLM, deeper semantic and intent features beyond traditional fixed features are extracted from the complex scenarios described by the prompts, thus addressing the problem of insufficient model generalization. More importantly, the system does not directly use the raw output of the LLM; instead, a downstream, targeted encoder relearns and adaptively transforms the LLM's prompts or hidden representations.
[0048] , in, Confirmation words indicating construction, Representing a large language model, This represents the multilayer perceptron algorithm. This represents the user's LLM coding characteristics.
[0049] Next, the characteristics of the graph As matrix, As and The input is fed into CrossAttention, which aligns the natural language task with the anomaly detection task.
[0050] , in, Feature representations that integrate semantic information express The matrix dimension corresponds to the number of columns. Corresponding to the softmax function, This represents the multilayer perceptron algorithm.
[0051] By short-term evolutionary features As Matrix, encoding features as and The matrix input cross-attention mechanism aligns and fuses natural language semantic tasks with graph anomaly detection tasks, outputting a feature representation with fused semantic information. . Figure 5This invention provides the effect of performing spatiotemporal graph walks on four real networks: Wikipedia, MOOC, Bitcoin-Alpha, and Bitcoin-OTC. The vertical axis represents the number of events in the data; the blue part is the original data, and the red part is the sampled data. Each data point is represented by K, which is thousands. The legend uses 70% of the data for testing, showing that after using spatiotemporal graph walks, the sampled data is significantly less than the original data. This invention only requires approximately 30% of the data from Wikipedia, 26% from MOOC, 50% from Bitcoin-Alpha, and 57% from Bitcoin-OTC to perform anomaly detection tasks.
[0052] Optionally, this also includes outputting multiple users based on their anomaly probabilities from highest to lowest, and outputting the number of users as required. Through MLP computation, this step aims to fuse long-term temporal evolution information and node attention information to obtain effective long-term temporal evolution node attention features.
[0053] , Let v represent the anomaly probability of all users, and let v be the anomaly probability at time i. Finally, output multiple nodes based on their anomaly probability from highest to lowest. The number of nodes output here depends on the actual situation. : , in, Indicates user Corresponding index The probability of an anomaly, The index corresponding to the user with the highest probability of abnormality among all users is selected, which is the Node of the prompt word.
[0054] The number of abnormal users output is a parameter set in advance by the user.
[0055] The calculated attention features are binary-classified using a neural network classifier to obtain anomaly scores for all users. Then, all users within the social area are traversed, and the users with the highest anomaly estimates are selected as the detected anomalous users. Here, "multiple users" refers to the number of anomalous users designated by the user. Figure 4 This paper demonstrates the method for calculating the anomaly estimates of users output by this invention. The user data in the illustration represents the aggregated attention features of users; the red portion represents known anomaly users, and the yellow portion represents users sampled by the spatiotemporal graph walk algorithm. The user data is input into a classifier using an MLP algorithm, and the output is the anomaly estimate for all users. The anomaly estimate value ranges from 0 to 1, with a higher value indicating a greater probability of being an anomalous user. The results are sorted in descending order of the anomaly estimate values. Output the estimated number of user anomalies in user settings, from largest to smallest. During the training phase, a clustering consistency loss function and... Figure 1 Consistency loss function optimizes model performance. Figure 6 This paper presents the anomaly detection performance of the present invention on four real networks: Wikipedia, MOOC, Bitcoin-Alpha, and Bitcoin-OTC. The vertical axis, "AUC(%)", represents the anomaly detection accuracy as a percentage, ranging from 0 to 100; a higher value indicates better anomaly detection. The horizontal axis, "Training Data Ratio", represents the proportion of data used in the experiment. 70%, 60%, 50%, 40%, and 30% of the original data were used for comparison to examine the algorithm's anomaly detection performance under limited sample conditions. The last 15% of the original data was used as test data. It can be seen that the anomaly detection performance of the present invention is superior to other methods on both datasets, and its anomaly detection performance decreases the least with decreasing input data, demonstrating its excellent performance in few-sample anomaly detection.
[0056] Optionally, this also includes using the output feature representation of the large language model during the training phase. Construct a loss function and optimize the model. The loss function includes the double consistency learning constraint loss function, that is... Figure 1 Consistency loss function and cluster consistency loss function. Figure 1 The consistency loss function aims to ensure that the local geometry of current LLM augmented node features remains consistent with their historical counterparts. The clustering consistency loss function aims to enforce distribution consistency by encouraging current LLM features and historical features to share similar cluster assignments in the latent space. Its purpose is to capture latent, more global, and time-series data features. Figure 7This invention provides the performance of a pre-trained model on the social network datasets Wikipedia and MOOC on the transaction network datasets Bitcoin-Alpha and Bitcoin-OTC for zero-shot anomaly detection. Wikipedia->Bitcoin-OTC indicates that the model weights from Wikipedia are used for zero-shot anomaly detection on the Bitcoin-OTC data. Zero-shot anomaly detection means that the algorithm is trained only on Wikipedia and MOOC, and directly detects anomalies without ever being trained on Bitcoin-Alpha and Bitcoin-OTC. It can be seen that the model weights using social network data perform better on the transaction network than other algorithms. This further demonstrates that the model trained using social network data in this invention can achieve excellent anomaly detection accuracy on the transaction network. The size of the test dataset is shown in Table 1.
[0057] Table 1. Size of the test dataset This invention reduces the training data requirement through spatiotemporal graph walking and combines it with a long-term memory to support incremental learning and rapid parameter updates, enabling the algorithm to adjust in real time based on the latest data and maintain strong timeliness and adaptability. Based on the double consistency learning constraint, only a small amount of known anomaly information is needed for model training and detection of other anomaly regions, eliminating the need for global anomaly information and reducing dependence on data integrity.
[0058] In summary, this invention relates to a method for detecting telecommunications fraud in financial transactions, capable of pinpointing anomalies to a very small area using limited known anomaly information. Locating anomalies within a small area not only improves prediction accuracy but also ensures minimal losses through early anomaly detection. Based on double consistency learning constraints, this algorithm can perform accurate anomaly detection with only limited anomaly data. It integrates graph data and general LLM knowledge through a double consistency loss function, stabilizing model training. This eliminates the need for costly and expensive anomaly information collection in reality, allowing for the training and detection of the origin tracing algorithm. Furthermore, the real-time updated memory model used in this invention possesses characteristics of real-world networks and the propagation of real-world anomaly information, making the origin tracing algorithm practically applicable. Finally, applying the real-time updated memory model and spatiotemporal graph walk method to anomaly detection in real-world networks demonstrates a strong ability to predict anomalies, providing a scientific basis for real-time rumor detection solutions on the internet.
[0059] The above-described embodiments are merely a few specific examples of the present invention. However, the embodiments of the present invention are not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A method of detecting financial transaction telecommunication fraud, characterized in that, Includes the following steps: Enter the user status of the target area where anomaly detection is required; Mapping the user status to a complex network wherein a set of users of a target region, a set of edges representing two users knowing each other in the social network, performing a spatio-temporal graph walk starting from the abnormal user to obtain information of other users related to the abnormal user; Statistically and update other user's operation sequence forms dynamic time characteristic ; Constructing each other user in complex network topology features of the interaction behavior between users ; Dynamic temporal features for the same user under one timestamp and topological features Aggregating to get data features ; Data features at all timestamps Perform fusion evolution to obtain short-term evolution features of the user ; Using prompt words to complex network The natural language input large language model is represented as a prompt word, the prompt word is relearned and adaptively converted by a downstream encoder, encoded features of a user are obtained, and short-term evolution features are obtained For guidance, correlation matching and weighted extraction are performed in the encoded features, and feature representations fused with semantic information are output ; The Non-linear transformation is performed, and the transformed result is mapped as a probability distribution through a normalization exponential transformation to obtain the anomaly probability of all users.
2. The financial transaction telecommunication fraud detection method of claim 1, wherein, The user status includes interactions between social network users.
3. The financial transaction telecommunication fraud detection method of claim 1, wherein, The spatiotemporal graph walk includes the following steps: Add abnormal users to a queue; Remove the abnormal user and other users associated with them from the queue, and access their connected neighbors; Add adjacent users to the visited set and queue, and do not sample users that have already been visited; The sampled data includes users, the time corresponding to the users, the feature vector corresponding to the users, the labels corresponding to the users, and fuzzy markings indicating whether the users are abnormal users; Output the sequence of users associated with the abnormal user.
4. The financial transaction telecommunication fraud detection method of claim 1, wherein, The statistics of the operation sequences of other users are updated to form dynamic time characteristics comprising the steps of: Generate a memory vector for each other user, with the initial memory vector being a zero matrix; When a new abnormal user and other users related to the new abnormal user need to be input, the memory vectors of the other users are updated, the memory vectors including operation sequences of the users and corresponding times , the operation sequences of the users including operation sequences of liking, commenting, and posting of the users Time The dimension of the memory vector is aligned by a single learnable nonlinear mapping function, and the previous memory vector is spliced to obtain a new feature matrix; The same operation of the same user is mapped to the same feature matrix by indexing, and the feature matrices of all other users form a dynamic time warping .
5. The financial transaction telecommunication fraud detection method of claim 1, wherein, The constitutive topological feature comprising the steps of: For the current one other user Limited sampling, sampling users who are friends or have interactive behaviors with each other , as its neighbor node , corresponding to the edge , this sampling only samples users less than the current time Sample and merge the corresponding user features Form a topological feature .
6. The financial transaction telecommunication fraud detection method of claim 1, wherein, The data features are obtained by taking the mean of the dynamic time features and the topological features for the same user.
7. The financial transaction telecommunication fraud detection method of claim 1, wherein, said short-term evolution characteristics of the user comprising the steps of: Data features at all timestamps Perform time series prediction to obtain new data features at the user node , input into the next calculation stage while updating the dynamic time features ; Obtaining data features of each user at a current timestamp and historical dynamic time features , and generating a time position encoding vector; splicing at the current timestamp with the time position encoding vector to obtain a feature matrix Splicing at the current timestamp, historical dynamic time features and the time position encoding vector to obtain a feature matrix and According to the feature matrix , and , the attention features of the user are obtained ; Will Input to the history memory queue, the queue outputs the characteristics of the historical data, and the current data characteristics and the characteristics of the history memory are spliced to obtain the short-term evolution characteristics of the user .
8. The financial transaction telecommunication fraud detection method of claim 7, wherein, The prompts include: a public introduction to the target area and a method for constructing user characteristics, user operation information, and timestamps of user status changes; By short-term evolutionary features As Matrix, encoding features as and The matrix input cross-attention mechanism aligns and fuses natural language semantic tasks with graph anomaly detection tasks, outputting a feature representation with fused semantic information. .
9. The financial transaction telecommunication fraud detection method of claim 1, wherein, It also includes outputting multiple users based on their abnormal probability from highest to lowest, and outputting the number of users as required.
10. The financial transaction telecommunication fraud detection method of claim 1, wherein, Also including large language models using outputted feature representations at training stage Constructing loss function, optimizing model.