Game user abnormal transaction behavior identification method and system

By constructing a personalized behavior baseline and dynamic transaction relationship diagram, combined with self-supervised learning and stream processing framework, the problem of insufficient recognition ability of complex transaction patterns in the existing technology is solved, and efficient identification and interpretation of abnormal transaction behaviors of game users is achieved, and the adaptability and interpretability of the model is enhanced.

CN120146856AActive Publication Date: 2025-06-13SHENZHEN DUI DUI TECH CO LTD

Patent Information

Application Number
CN202510632000.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing game transaction recognition methods are difficult to cope with complex and changeable trading models, static models cannot adaptively adjust, lack effective identification capabilities for new abnormal trading forms, and the model is poorly interpretable, making it difficult to clearly explain the basis for abnormal judgment to players and operators.

Method used

By obtaining multimodal data of users' transactions, in-game behavior, social relationship data and device environment, a personalized behavior baseline and dynamic transaction relationship diagram are constructed, and a loss function is constructed using self-supervised learning and comparison learning methods, real-time transaction data processing is combined with a stream processing framework, comprehensive risk scores are calculated, and model decisions are explained through SHAP value, and abnormal transaction interpretation labels and evidence links are generated.

Benefits of technology

It realizes efficient identification and interpretation of abnormal trading behaviors of game users, can dynamically adjust detection strategies and model parameters, adapt to changes in the game environment, reduces the workload and cost of manual labeling, enhances the depth of understanding and generalization capabilities of the model, and improves the trust of players and operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146856A_ABST
    Figure CN120146856A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of game transaction management, and provides a game user abnormal transaction behavior identification method and system, and the method comprises the following steps: S1, obtaining the transaction of a user, the behavior in a game, social relation data and the multi-modal data of an equipment environment, constructing a personalized behavior baseline based on the historical data of the user, generating a user normal behavior sequence; and S2, regarding the user as graph nodes, transactions, behaviors, social contact and equipment association as edges, constructing a dynamic transaction relation graph, and analyzing the updated transaction relation graph according to an edge weight calculation result. A dynamic transaction model is constructed by using a dynamic transaction construction module, and a reinforcement learning strategy module is combined, and a dynamic adjustment detection strategy and model parameters are fed back in real time according to a game environment, so that changes of economic system adjustment, new transaction mode occurrence and the like in a game can be effectively coped with, and efficient recognition capability on abnormal transaction behaviors is continuously kept; the defect of poor adaptability of a traditional static model is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of game transaction management, and particularly to a method and system for identifying abnormal transaction behaviors of game users. Background Art

[0002] At present, with the booming development of the game industry, virtual item trading is frequent and player interaction is active. The game economic system is becoming increasingly complex. The transaction behaviors of game users cover the transfer of various virtual assets such as virtual currency, props, and equipment, and the transaction forms are also becoming more and more diverse. However, driven by interests, some users damage the game economic balance through abnormal transaction behaviors such as malicious coin brushing, account sharing, and false transactions, which harm the experience of other players and the interests of game operators.

[0003] Traditional transaction identification methods rely on manually set rules, which are difficult to cope with complex and changeable transaction patterns, and it is easy to miss detections due to untimely rule updates. Although machine learning methods have been applied, most of them are static models and cannot be adaptively adjusted with the dynamic changes of the game environment, lacking the ability to effectively identify newly emerging abnormal transaction forms. In addition, existing methods perform poorly in terms of model interpretability, making it difficult to clearly explain the basis for abnormal determination to players and operators, leading to trust crises. Summary of the Invention

[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method and system for identifying abnormal transaction behaviors of game users to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides a method for identifying abnormal transaction behaviors of game users, including the following steps:

[0006] S1. Obtain multi-modal data of the user's transactions, in-game behaviors, social relationship data, and device environment, and construct a personalized behavior baseline based on the user's historical data to generate a sequence of the user's normal behaviors;

[0007] S2. Regard the user as a graph node, and use transactions, behaviors, social connections, and device associations as edges to construct a dynamic transaction relationship graph, and analyze the updated transaction relationship graph according to the edge weight calculation results;

[0008] S3. Use unlabeled game log data for self-supervised training, adopt a contrastive learning method to construct a loss function, and calculate virtual economic balance indicators according to the in-game item drop mechanism, currency inflation rate, and transaction tax rules;

[0009] S4. Use a stream processing framework to process real-time transaction data at the millisecond level, calculate a comprehensive risk score, and divide risk levels according to the risk score;

[0010] S5. Use SHAP values to explain model decisions, generate abnormal transaction explanation labels, form an immutable evidence chain for abnormal transactions, and collect and annotate data through human-machine collaboration to reverse-optimize model parameters and risk scores;

[0011] S6. Based on the feedback of each abnormal transaction determination process and optimization result, generate dynamic detection strategies through a policy network, and continuously interact with the game environment to optimize the parameters of the policy network.

[0012] Preferably, in the step S1, generating the normal behavior sequence of users includes the following steps:

[0013] S11. Collect transaction data between users, in-game behavior data, social relationship data, and device environment data information, and clean and normalize the data;

[0014] S12. Divide time into time windows of fixed length. Within each time window, calculate the statistics of each feature. The formula is: , where is the mean value, is the standard deviation, is the number of transactions within the time window, is the th transaction amount;

[0015] S13. Use the feature statistics of each time window as the input sequence and output the predicted feature values. The formula is: , where is the output predicted feature value, is the weight matrix of the fully connected layer, is the hidden state obtained after the last time step, is the bias vector;

[0016] S14. According to the calculation results of the mean value and standard deviation, establish a stationary autoregressive model and a moving average model. The formula is: , where is the sequence of the autoregressive part after stationary transformation, is the dependent variable of the moving average part, is the autoregressive order, is the autoregressive coefficient, is the value of the stationary time series at time, is the moving average order, is the moving average coefficient, and are the values of the white noise sequence at and time respectively;

[0017] S15. Combine the established stationary autoregressive model and moving average model to establish an ARIMA model. The formula is: , where in the formula, is the lag operator, and , is the order of differencing, is the autoregressive polynomial, and , where are the autoregressive coefficients, is the order of autoregression, is the moving average polynomial, , where are the moving average coefficients;

[0018] S16. Use the estimated ARIMA model with parameters to predict future values and obtain the finally predicted eigenvalue .

[0019] Preferably, in the step S2, the analysis of the updated transaction relationship graph includes the following steps:

[0020] S21. Abstract game users as nodes to construct a node set, and abstract the transaction, social, and device relationships of users as edges to form an edge set, and construct a dynamic transaction relationship graph. The formula is: , where in the formula, is the dynamic transaction relationship graph, is the node set, is the edge set;

[0021] S22. Combine the results of multi-modal data collection, assign an initial feature vector to each node, and calculate the weight of the edge, which reflects the closeness of the relationship between users. The formula is: , where in the formula, is the weight of the edge. The larger the value, the closer the relationship between users, is the transaction amount, is the transaction frequency, is the social intimacy between users, and the value range is [0 - 1], , and are the corresponding weights respectively, and are the normalized transaction amount and frequency respectively;

[0022] S23. Aggregate the node features in the neighborhood and update the feature vector of the node. The formula is: , where in the formula, is the feature vector of node at the th layer, is the set of neighborhood nodes of the node , is the weight matrix of the th layer, is the activation function, is to average the features of neighborhood nodes;

[0023] S24. Process the updated transaction relationship graph to identify high-density transaction clusters or one-sided star network structures in the graph. The formula is , where is the modularity value, is the total number of edges in the transaction relationship graph, is the node and the node The edge weight between them, and are the nodes and the node Degree of, is to judge whether the node and the node belong to the same community. If they belong, it is 1, otherwise it is 0. When it is detected that the modularity value suddenly changes or a subgraph structure that does not conform to the normal transaction pattern appears, it is determined as an abnormal subgraph, and there may be abnormalities in the transactions between the corresponding users.

[0024] Preferably, in the step S3, the self-supervised training and abnormal pattern recognition include the following steps:

[0025] S31. Collect game log data, including but not limited to unlabeled data of normal transactions, behavior trajectories, and social relationships, and construct an unlabeled data set;

[0026] S32. Perform data augmentation operations on the original feature vectors of each user in the unlabeled data set to generate augmented feature vectors;

[0027] S33. Based on the idea of contrastive learning, construct a loss function to enable the model to learn the similarity of different augmented views of the same user and the difference of feature vectors of different users. The formula is: , where Feature vector value of the cross-loss function, is the original feature vector, is the augmented feature vector, is the feature vector of other users, is the temperature parameter that controls the learning difficulty, is the cosine similarity function that measures the similarity degree of two feature vectors;

[0028] S34. Train the model using the constructed loss function. During the training process, input the original user feature vectors and augmented feature vectors in the unlabeled dataset into the model, calculate the loss, and update the model parameters through backpropagation to optimize the weights of the model, enabling the model to learn general trading patterns and behavioral rules from the unlabeled data;

[0029] S35. After training is completed, introduce a small amount of labeled abnormal trading data to fine-tune the pre-trained model. Input the labeled data into the model and further optimize the model parameters through the cross-loss function to make the model adapt to the specific requirements of the game abnormal trading scenario and output the abnormal probability score. The formula is: , where is the total cross-entropy loss of the entire labeling technology, is the number of samples in the labeled abnormal trading data, is the -th true label of the sample, is the abnormal probability score output by the model for the -th sample , is the -th input trading data feature;

[0030] S36. According to the abnormal probability score output by the model, set a threshold . When , determine that the user's trading behavior is abnormal. When , determine that the user's trading behavior is normal.

[0031] Preferably, in step S32, generating the augmented feature vector includes the following steps:

[0032] S321. Randomly replace some information in the trading item type with special markers, forcing the model to learn to restore complete features from incomplete information;

[0033] S322. Add a random time offset to the trading timestamp to simulate small fluctuations in the time series and enhance the model's robustness to time changes;

[0034] S323. Randomly sample some subgraphs from the user's social relationship graph, retain the key social links, and remove the honor information to make the model focus on the core social features.

[0035] Preferably, in step S4, processing the trading data and dividing the risk levels includes the following steps:

[0036] S41. Intervene in game transaction records in real time through a stream processing framework, including but not limited to the timestamp of each transaction, the IDs of the two parties to the transaction, the transaction amount, and the transaction item information, and generate real-time feature vectors.

[0037] S42. Use edge computing nodes to preprocess transaction data locally in the server cluster, extract lightweight graph features related to the device environment, including but not limited to device fingerprints, IP addresses, and whether an emulator is in use, and generate unique identifiers.

[0038] S43. Compare the real-time feature vectors with the dynamic behavior baseline and calculate the deviation. The formula is: , where is the real-time feature vector is the th eigenvalue of , is the th eigenvalue of the behavior baseline vector , is the feature dimension. The larger the value of

[0039] S44. Incorporate real-time transaction data into the dynamic transaction relationship graph, update the graph structure and node features, and analyze the relationship graph to output the anomaly score of each transaction behavior. The formula is: , where is the anomaly score of node , is the dimension of the embedding vector, and are the th components of the embedding vector of node and the centroid vector of the abnormal transaction cluster respectively;

[0040] S45. Calculate the comprehensive risk score based on the anomaly score, combined with the dynamic behavior baseline deviation and the transaction anomaly score. The formula is: , where and and are the corresponding weight coefficients respectively, , and divide the risk threshold. According to the risk score, divide the risk level. When , it is classified as low risk. When , it is classified as medium risk. When , it is classified as high risk. When and and are the set low, medium, and high risk thresholds respectively.

[0041] Preferably, in the step S5, generating the abnormal transaction evidence chain, the reverse optimization model parameters and the risk score include the following steps:

[0042] S51. Calculate the SHAP value to quantify the contribution degree of each feature to the model prediction result. The formula is: , where in the formula, is the SHAP value of the feature , is a subset in the feature set , is the number of elements in the subset , is the total number of features, is the predicted value of the model when only using the features in the subset , is the change in the predicted value of the model after adding the feature to the subset ;

[0043] S52. Screen out the features with larger absolute values of SHAP values, and combine with the preset interpretation template to convert the key features and their SHAP values into natural language;

[0044] S53. Collect evidence data, and perform a hash operation on the collected evidence data to generate a unique transaction hash value. The formula is: , where in the formula, is the transaction hash value, is data concatenation, , , , , and are the transaction timestamp, transaction amount, transaction item information, device fingerprint, interpretation label of the SHAP value, and the output abnormal probability score respectively. Then record the transaction hash value, original evidence data, timestamp, transaction amount, transaction item information, device fingerprint, interpretation label of the SHAP value, and the output abnormal probability score into the blockchain distributed ledger to ensure the authenticity and permanent storage of the evidence chain;

[0045] S54. The system displays the preliminary determination result of the abnormal transaction, the SHAP value interpretation label, and the original evidence data to the manual review personnel for manual determination and annotation, and retrieves the manually annotated data;

[0046] S55. Divide the retrieved annotated data into a training set and a test set, and adopt the optimization algorithm of gradient descent to perform backpropagation on the model parameters. The formula is: , where in the formula, is the eigenvector value of the cross-loss function after backpropagation, is the sample The true label, is the predicted probability of the model for the sample .

[0047] Preferably, in the step S6, optimizing the policy network parameters includes the following steps:

[0048] S61. Integrate the economic state, user behavior data, and model recognition results of the current game into a state vector. The formula is: , where in the formula, is the integrated state vector, is the game economic indicator vector, including but not limited to the game currency inflation rate and the deviation degree of rare equipment prices, is the deviation amount of the user behavior baseline, is the historical risk level sequence, is the abnormal probability score output by the model;

[0049] S62. Define executable detection policy actions and design a reward mechanism according to the abnormal transaction determination and optimization results. The formula is: , where in the formula, is the reward value, is the change rate of the abnormal transaction recognition accuracy, is the improvement value of the game economic indicator stability, is the number of false alarms, , and are the corresponding weights respectively;

[0050] S63. Construct a policy network, output the probability distribution of each action, and optimize the policy network parameters according to the dynamic policy;

[0051] S64. Repeat steps S61 - S63, continuously interact with the game environment, accumulate experience data and update the policy network until the loss function converges or the reward value reaches a stable state, and then output the final dynamic detection policy.

[0052] Preferably, in the step S63, optimizing the policy network parameters includes the following steps:

[0053] S631. Use a deep neural network to construct a policy network, input the state vector, and output the probability distribution of the action. The formula is: , where in the formula, is the policy network takes the action under the given state with the probability of is the parameter set of the policy network, and are the network weight matrices, and is the paranoia vector;

[0054] S632. Randomly initialize the parameters of the policy network, and set the learning rate and discount factor parameters;

[0055] S633. At each time step, sample an action according to the current state through the policy network, and perform the corresponding detection policy adjustment;

[0056] S634. After performing the sampled action, observe the change of the game environment, and obtain the new state and the reward value;

[0057] S635. Adopt the temporal difference learning algorithm to define the objective function, and the formula is: , where is the state-action value function when performing action under the policy at state , is the expectation of the internal random variable, is the reward value obtained at time is the discount factor, and its value range is [0 - 1]. The closer its value is to 0, the more it pays attention to the current reward. The closer its value is to 1, the more it pays attention to the future reward, is the maximum state-action value when performing the next action under the policy at the next state ;

[0058] S636. Optimize the network policy by minimizing the loss function, and the formula is: , where is the loss function value, is the logarithm of the probability of the policy network taking action at the given state ;

[0059] S637. Update the network policy using stochastic gradient descent, and the formula is: , where is the learning rate, is the loss function with respect to the parameter gradient.

[0060] A game user abnormal transaction behavior recognition system, which is applied to any one of the above-mentioned game user abnormal transaction behavior recognition methods, and includes:

[0061] A data acquisition and processing module, which is used to acquire multi-source data and preprocess the data;

[0062] A dynamic transaction construction module for constructing a user's dynamic transaction model based on the processed data;

[0063] A self-supervised learning module for mining the internal structure and laws of data without manual annotation;

[0064] A risk assessment and division module for risk assessment of user transaction behaviors based on the output feature representation;

[0065] A traceability and hierarchical processing module for performing traceability analysis on abnormal transactions determined to be of high risk;

[0066] A reinforcement learning strategy module for continuously optimizing the detection strategy according to the system feedback, and dynamically adjusting the model parameters and detection rules.

[0067] The beneficial effects of a method and system for identifying abnormal transaction behaviors of game users provided by the present invention are as follows:

[0068] 1. By using the dynamic transaction construction module to construct a dynamic transaction model and combining it with the reinforcement learning strategy module, and dynamically adjusting the detection strategy and model parameters according to the real-time feedback of the game environment, it can effectively cope with changes such as the adjustment of the in-game economic system and the emergence of new transaction modes, continuously maintain the high-efficiency identification ability of abnormal transaction behaviors, and overcome the disadvantages of poor adaptability of traditional static models.

[0069] 2. Through the self-supervised learning module with the help of self-supervised learning technology, it can mine the internal laws of data without a large amount of manually annotated data, learn normal and abnormal transaction modes, greatly reduce the workload and cost of manual annotation, and at the same time improve the understanding depth and generalization ability of the model for transaction modes, enabling the model to still operate well under limited data.

[0070] 3. Through the traceability and hierarchical processing module, traceability analysis and hierarchical disposal are carried out for transactions of different risk levels, realizing the precise strike of abnormal transaction behaviors, protecting the game economic order, reasonably controlling the processing intensity, avoiding excessive interference with normal transactions, and realizing the interpretability of model decisions through technologies such as SHAP values, generating clear and easy-to-understand abnormal transaction explanation labels, clearly showing the reasons for abnormal determination and key influencing factors to players and operators, and enhancing players' trust in the fairness of game operation. Description of the Drawings

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 Schematic diagram of the step flow of a method and system for identifying abnormal trading behaviors of game users provided by this application;

[0073] Figure 2 Schematic diagram of the system modules of a method and system for identifying abnormal trading behaviors of game users provided by this application. Detailed implementation manners

[0074] The following further describes in detail the specific implementation manners of the present invention in conjunction with the specification drawings and embodiments. The following embodiments are only used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0075] As Figure 1 - Figure 2 shown, this embodiment proposes a method for identifying abnormal trading behaviors of game users, including the following steps:

[0076] S1. Obtain multi-modal data of the user's transactions, in-game behaviors, social relationship data, and device environment, and construct a personalized behavior baseline based on the user's historical data to generate a sequence of normal user behaviors;

[0077] S2. Regard the user as a graph node, and use transactions, behaviors, social interactions, and device associations as edges to construct a dynamic trading relationship graph. Analyze the updated trading relationship graph according to the edge weight calculation results;

[0078] S3. Use unlabeled game log data for self-supervised training, adopt a contrastive learning method to construct a loss function, and calculate the virtual economy balance index according to the in-game item drop mechanism, currency inflation rate, and transaction tax rules;

[0079] S4. Perform millisecond-level processing on real-time transaction data through a stream processing framework, calculate a comprehensive risk score, and divide the risk level according to the risk score;

[0080] S5. Use SHAP values to explain model decisions, generate abnormal trading explanation labels, form an immutable evidence chain of abnormal transactions, and collect labeled data through human-computer collaboration to reversely optimize model parameters and risk scores;

[0081] S6. Based on the feedback of each abnormal transaction determination process and optimization processing result, generate a dynamic detection strategy through a policy network, and continuously interact with the game environment to optimize the policy network parameters.

[0082] In this embodiment, in step S1, generating a sequence of normal user behaviors includes the following steps:

[0083] S11. Collect transaction data, in-game behavior data, social relationship data, and device environment data information between users, and perform data cleaning and normalization processing;

[0084] S12. Divide the time into time windows of fixed length. Within each time window, calculate the statistics of each feature. The formula is: , where in the formula is the mean, is the standard deviation, is the number of transactions within the time window, is the th transaction amount;

[0085] S13. Take the feature statistics of each time window as the input sequence and output the predicted feature values. The formula is: , where in the formula is the output predicted feature value, is the weight matrix of the fully connected layer, is the hidden state obtained after the last time step, is the bias vector;

[0086] S14. According to the calculation results of the mean and standard deviation, establish a stationary autoregressive model and a moving average model for the sequence. The formula is: , where in the formula is the sequence after stationary autoregressive part, is the dependent variable of the moving average part, is the autoregressive order, is the autoregressive coefficient, is the value of the stationary time series at time, is the moving average order, is the moving average coefficient, and are the values of the white noise sequence at and time respectively;

[0087] S15. Combine the established stationary autoregressive model and moving average model to establish an ARIMA model. The formula is: , where in the formula is the lag operator and , is the difference order, is the autoregressive polynomial and , where is the autoregressive coefficient, is the autoregressive order, is the moving average polynomial, , where is the moving average coefficient;

[0088] S16. Use the ARIMA model with estimated parameters to predict future values and obtain the finally predicted eigenvalue. .

[0089] Specifically, the data acquisition and processing module comprehensively obtains multi-modal data such as user transactions, in-game behaviors, social relationships, and device environments, and through in-depth analysis and feature extraction, fully explores the potential associations between data. Compared with traditional methods that only rely on single or a small number of data dimensions, it can more accurately capture the characteristics of abnormal transaction behaviors and significantly reduce the false detection and missed detection rates.

[0090] In this embodiment, in step S2, the analysis of the updated transaction relationship graph includes the following steps:

[0091] S21. Abstract game users as nodes to construct a node set, and abstract the transaction, social, and device relationships of users as edges to form an edge set, and construct a dynamic transaction relationship graph. The formula is: , where is the dynamic transaction relationship graph, is the node set, is the edge set;

[0092] S22. Combine the results of multi-modal data acquisition, assign an initial feature vector to each node, and calculate the weight of the edge, which reflects the closeness of the relationship between users. The formula is: , where is the weight of the edge, and the larger the value, the closer the relationship between users, is the transaction amount, is the transaction frequency, is the social intimacy between users, with a value range of [0 - 1], , and are the corresponding weights respectively, and are the normalized transaction amount and frequency respectively;

[0093] S23. Aggregate the neighborhood features of the node and update the feature vector of the node. The formula is: , where is the feature vector of node at the th layer, is the set of neighborhood nodes of node , is the weight matrix at the th layer, is the activation function, is to take the average of the neighborhood node features;

[0094] S24. Process the updated transaction relationship graph to identify high-density transaction clusters or one-sided star network structures in the graph. The formula is , where in the formula, is the modularity value, is the total number of edges in the transaction relationship graph, is the node and the node The edge weight between them, and are respectively the degrees of the node and the node , is a function to judge whether the node and the node belong to the same community. If they belong, it is 1, otherwise it is 0. When it is detected that the modularity value suddenly changes or there is a subgraph structure that does not conform to the normal transaction pattern, it is determined as an abnormal subgraph, and there may be abnormalities in the transactions between the corresponding users.

[0095] In this embodiment, in step S3, self-supervised training and abnormal pattern recognition include the following steps:

[0096] S31. Collect game log data, including but not limited to unlabeled data of normal transactions, behavior trajectories, and social relationships, and construct an unlabeled data set;

[0097] S32. Perform data augmentation operations on the original feature vectors of each user in the unlabeled data set to generate augmented feature vectors;

[0098] S33. Based on the idea of contrastive learning, construct a loss function so that the model learns the similarity of different augmented views of the same user and the difference of feature vectors of different users. The formula is: , where in the formula, is the feature vector value of the cross-loss function, is the original feature vector, is the augmented feature vector, is the feature vector of other users, is the temperature parameter to control the learning difficulty, is the cosine similarity function to measure the similarity degree of two feature vectors;

[0099] S34. Use the constructed loss function to train the model. During the training process, input the original feature vectors and augmented feature vectors of the users in the unlabeled data set into the model, calculate the loss and update the model parameters through backpropagation to optimize the weights of the model, so that the model can learn the general transaction patterns and behavior rules from the unlabeled data;

[0100] S35. After the training is completed, a small amount of labeled abnormal transaction data is introduced to fine-tune the pre-trained model. The labeled data is input into the model, and the model parameters are further optimized through the cross-loss function to make the model adapt to the specific requirements of the game abnormal transaction scenario, and an abnormal probability score is output. The formula is as follows: , where is the total cross-entropy loss of the entire labeling technology, is the number of samples in the labeled abnormal transaction data, is the true label of the th sample, is the abnormal probability score output by the model for the th sample is the th input transaction data feature;

[0101] S36. According to the abnormal probability score output by the model, a threshold is set. When , it is determined that the user's transaction behavior is abnormal. When , it is determined that the user's transaction behavior is normal.

[0102] In this embodiment, in step S32, generating the augmented feature vector includes the following steps:

[0103] S321. Randomly replace some information in the transaction item type with special markers to force the model to learn to recover the complete features from the incomplete information;

[0104] S322. Add a random time offset to the transaction timestamp to simulate the minor fluctuations in the time series and enhance the model's robustness to time changes;

[0105] S323. Randomly sample some subgraphs from the user's social relationship graph, retain the key social links, and remove the honor information to make the model focus on the core social features.

[0106] Specifically, through the self-supervised learning module with the help of self-supervised learning technology, the internal laws of the data can be mined without a large amount of manually labeled data, the normal and abnormal transaction patterns can be learned, the manual labeling workload and cost can be greatly reduced, and at the same time, the model's understanding depth and generalization ability of the transaction pattern can be improved, so that the model can still operate well under limited data.

[0107] In this embodiment, in step S4, processing the transaction data and dividing the risk levels includes the following steps:

[0108] S41. Intervene in game transaction records in real time through a stream processing framework, including but not limited to the timestamp of each transaction, the IDs of the two parties to the transaction, the transaction amount, and the transaction item information, and generate real-time feature vectors;

[0109] S42. Use edge computing nodes to preprocess transaction data locally in the server cluster, extract lightweight graph features related to the device environment, including but not limited to device fingerprints, IP addresses, and whether there is an emulator in use, and generate unique identifiers;

[0110] S43. Compare the real-time feature vector with the dynamic behavior baseline and calculate the deviation. The formula is: , where is the real-time feature vector is the th eigenvalue of is the behavior baseline vector is the th eigenvalue of is the feature dimension. When has a larger value, it indicates a higher degree of deviation of the current transaction behavior from the normal mode;

[0111] S44. Incorporate real-time transaction data into the dynamic transaction relationship graph, update the graph structure and node features, and analyze the relationship graph to output the anomaly score of each transaction behavior. The formula is: , where is the anomaly score of node , is the dimension of the embedding vector, and are the th components of the embedding vector of node and the centroid vector of the abnormal transaction cluster respectively;

[0112] S45. Calculate the comprehensive risk score based on the anomaly score, combined with the dynamic behavior baseline deviation and the transaction anomaly score. The formula is: , where and and are the corresponding weight coefficients respectively, , and divide the risk threshold. According to the risk score, divide the risk level. When , it is classified as low risk. When , it is classified as medium risk. When , it is classified as high risk, when and and are the set low, medium, and high risk thresholds respectively.

[0113] In this embodiment, in step S5, generating the abnormal transaction evidence chain, the reverse optimization model parameters, and the risk score includes the following steps:

[0114] S51. Calculate the SHAP value to quantify the contribution of each feature to the model prediction result. The formula is: , where is the SHAP value of feature , is a subset of the feature set , is the number of elements in the subset , is the total number of features, is the predicted value of the model when only using the features in the subset , is the change in the model prediction value after adding the feature to the subset ;

[0115] S52. Screen out the features with larger absolute values of SHAP values, and combine with the preset interpretation template to convert the key features and their SHAP values into natural language;

[0116] S53. Collect evidence data, perform a hash operation on the collected evidence data to generate a unique transaction hash value. The formula is: , where is the transaction hash value, is data concatenation, , , , , and are the transaction timestamp, transaction amount, transaction item information, device fingerprint, the interpretation label of the SHAP value, and the output abnormal probability score respectively. Then record the transaction hash value, the original evidence data, the timestamp, the transaction amount, the transaction item information, the device fingerprint, the interpretation label of the SHAP value, and the output abnormal probability score into the blockchain distributed ledger to ensure the authenticity and permanent storage of the evidence chain;

[0117] S54. The system displays the preliminary determination result of the abnormal transaction, the SHAP value interpretation label, and the original evidence data to the manual review personnel for manual determination and annotation, and retrieves the manually annotated data;

[0118] S55. Divide the retrieved annotated data into a training set and a test set, and use the gradient descent optimization algorithm to perform backpropagation on the model parameters. The formula is: , where is the feature vector value of the cross-loss function after backpropagation, is the sample The true label is the predicted probability of the model for the sample .

[0119] Specifically, through the traceability and stratification processing module, traceability analysis and stratified disposal are carried out for transactions with different risk levels, so as to accurately strike abnormal transaction behaviors, ensure the game economic order, reasonably control the processing intensity, avoid excessive interference with normal transactions, and realize the interpretability of model decisions through technologies such as SHAP values, generate clear and easy-to-understand abnormal transaction explanation labels, clearly show the reasons for abnormal determination and key influencing factors to players and operators, and enhance players' trust in the fairness of game operation.

[0120] In this embodiment, in step S6, optimizing the policy network parameters includes the following steps:

[0121] S61. Integrate the economic state of the current game, user behavior data, and model recognition results into a state vector. The formula is: , where is the integrated state vector, is the game economic index vector, including but not limited to the game currency inflation rate and the deviation degree of rare equipment prices, is the baseline deviation of user behavior, is the historical risk level sequence, is the abnormal probability score output by the model;

[0122] S62. Define executable detection policy actions and design a reward mechanism according to the abnormal transaction determination and optimization results. The formula is: , where is the reward value, is the change rate of the accuracy of abnormal transaction identification, is the improvement value of the stability of game economic indicators, is the number of false alarms, , and are the corresponding weights respectively;

[0123] S63. Construct a policy network, output the probability distribution of each action, and optimize the policy network parameters according to the dynamic policy;

[0124] S64. Repeat steps S61 - S63, continuously interact with the game environment, accumulate experience data and update the policy network until the loss function converges or the reward value reaches a stable state, and then output the final dynamic detection policy.

[0125] In this embodiment, in step S63, optimizing the policy network parameters includes the following steps:

[0126] S631. Construct a policy network using a deep neural network, input the state vector, and output the probability distribution of actions. The formula is: , where is the policy network is the probability of taking action under the given state , is the parameter set of the policy network, and are the network weight matrices, and are the bias vectors;

[0127] S632. Randomly initialize the parameters of the policy network and set the learning rate and discount factor parameters;

[0128] S633. At each time step, sample an action according to the current state through the policy network and perform the corresponding detection policy adjustment;

[0129] S634. After performing the sampled action, observe the changes in the game environment and obtain the new state and the reward value;

[0130] S635. Adopt the temporal difference learning algorithm to define the objective function. The formula is: , where is the state-action value function of taking action under the policy at state , is to take the expectation of the internal random variable, is the reward value obtained at time, is the discount factor, and its value range is [0 - 1]. The closer its value is to 0, the more it focuses on the current reward. The closer its value is to 1, the more it focuses on the future reward, is the maximum state-action value of taking the next action at the next state under the policy ;

[0131] S636. Optimize the network policy by minimizing the loss function. The formula is: , where is the loss function value, is the policy network taking action under the given state the logarithm of the probability;

[0132] S637. Use stochastic gradient descent to update the network policy. The formula is: , where is the learning rate, is the loss function for the parameter gradient of

[0133] Specifically, by using the dynamic transaction construction module to construct a dynamic transaction model, combining with the reinforcement learning strategy module, and dynamically adjusting the detection strategy and model parameters in real-time according to the feedback of the game environment, it can effectively cope with changes such as the adjustment of the in-game economic system and the emergence of new transaction modes, continuously maintain the high-efficiency recognition ability of abnormal transaction behaviors, and overcome the disadvantages of poor adaptability of traditional static models.

[0134] A system for identifying abnormal transaction behaviors of game users, which is applied to a method for identifying abnormal transaction behaviors of game users in any one of the above, includes:

[0135] The data acquisition and processing module is used to collect multi-source data and preprocess the data;

[0136] The dynamic transaction construction module is used to construct a dynamic transaction model of the user based on the processed data;

[0137] The self-supervised learning module is used to mine the internal structure and rules of the data without manual annotation;

[0138] The risk assessment and division module is used to conduct risk assessment on the user's transaction behavior according to the output feature representation;

[0139] The traceability and hierarchical processing module is used to conduct traceability analysis on the abnormal transactions determined to be of high risk;

[0140] The reinforcement learning strategy module is used to continuously optimize the detection strategy according to the system feedback, and dynamically adjust the model parameters and detection rules.

[0141] The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Although the present invention has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications or equivalent replacements of the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and should all be covered by the scope of the claims of the present invention.

Claims

1. A method for identifying abnormal transaction behavior of game users, characterized in that: The following steps are involved: S1. Obtain multimodal data on users’ transactions, in-game behaviors, social relationships, and device environments, and build a personalized behavior baseline based on user historical data to generate a normal user behavior sequence. S2. Consider users as graph nodes, transactions, behaviors, social interactions, and device associations as edges, build a dynamic transaction relationship graph, and analyze the updated transaction relationship graph based on the edge weight calculation results; S3. Use unlabeled game log data for self-supervised training, use contrastive learning to construct the loss function, and calculate the virtual economic balance index based on the in-game item drop mechanism, currency inflation rate, and transaction tax rules; S4, processes real-time transaction data in milliseconds through the stream processing framework, calculates the comprehensive risk score, and divides the risk level according to the risk score; S5. Use SHAP values ​​to explain model decisions, generate abnormal transaction explanation labels, form an unalterable abnormal transaction evidence chain, and collect labeled data through human-machine collaboration to reversely optimize model parameters and risk scores; S6. Based on the feedback of each abnormal transaction determination process and optimization processing results, a dynamic detection strategy is generated through the strategy network, and it continuously interacts with the game environment to optimize the strategy network parameters.

2. A method for identifying abnormal transaction behavior of game users according to claim 1, characterized in that: In step S1, generating a normal user behavior sequence includes the following steps: S11. Collect transaction data between users, in-game behavior data, social relationship data, and device environment data information, and clean and normalize the data; S12. Divide the time into time windows of fixed length, and calculate the statistics of each feature in each time window. The formula is: , where is the mean, is the standard deviation, is the number of transactions in the time window, For the The transaction fee of the times; S13. Take the characteristic statistics of each time window as the input sequence and output the predicted characteristic value. The formula is: , where is the feature value of the output prediction, is the weight matrix of the fully connected layer, is the hidden state obtained after the last time step, is the bias vector; S14. According to the calculation results of mean and standard deviation, establish the autoregressive model and moving average model of the stabilized sequence. The formula is: , where is the sequence after the natural regression part is stabilized, is the dependent variable of the moving average part, is the autoregressive order, is the autoregression coefficient, The time series after stationarization is The value of the moment, is the moving average order, is the moving average coefficient, and are white noise sequences in and The value of the moment; S15. Combine the established stable sequence autoregressive model and moving average model to establish the ARIMA model. The formula is: , where is a lag operator, and , is the difference order, is an autoregressive polynomial, and ,in, is the autoregression coefficient, is the autoregressive order, is the moving average polynomial, ,in, is the moving average coefficient; S16. Use the ARIMA model with estimated parameters to predict future values ​​and obtain the final predicted eigenvalues .

3. A method for identifying abnormal transaction behavior of game users according to claim 1, characterized in that: In step S2, analyzing the updated transaction relationship graph includes the following steps: S21. Game users are abstracted as nodes, and node sets are constructed. User transactions, social interactions, and device relationships are abstracted as edges to form edge sets. A dynamic transaction relationship graph is constructed. The formula is: , where It is a dynamic transaction relationship diagram. is a node set, For the edge collection; S22. Combined with the multimodal data collection results, an initial feature vector is assigned to each node, and the weight of the edge is calculated to reflect the closeness of the relationship between users. The formula is: , where is the edge weight. The larger the value, the closer the relationship between users. For transaction funds, is the transaction frequency, is the social intimacy between users, with a value range of [0-1], , and are the corresponding weights, and are the normalized transaction amount and frequency respectively; S23, perform neighborhood aggregation on node features and update the feature vector of the node. The formula is: , where For Node In the The feature vector of the layer, For Node The collection of neighboring nodes of For the The weight matrix of the layer, is the activation function, To average the features of the neighborhood nodes; S24. Process the updated transaction relationship graph to identify high-density transaction clusters or one-side star network structures in the graph. The formula is: , where is the modularity value, is the total number of edges in the transaction relationship graph, For Node and nodes The edge weights between and Node and nodes The degree, To judge the node and nodes A function of whether they belong to the same community. If they do, it is 1, otherwise it is 0. When a sudden change in the modularity value or a subgraph structure that does not conform to the normal transaction pattern is detected, it is determined to be an abnormal subgraph, and the transactions between the corresponding users may be abnormal.

4. A method for identifying abnormal transaction behavior of game users according to claim 1, characterized in that: In step S3, the self-supervised training and abnormal pattern recognition includes the following steps: S31. Collect game log data, including but not limited to unlabeled data of normal transactions, behavior trajectories, and social relationships, and construct an unlabeled data set; S32, performing a data enhancement operation on the original feature vector of each user in the unlabeled data set to generate an augmented feature vector; S33. Based on the idea of ​​contrastive learning, a loss function is constructed so that the model can learn the similarities between different augmented views of the same user and the differences between feature vectors of different users. The formula is: , where The eigenvector value of the cross loss function, is the original feature vector, is the augmented eigenvector, are feature vectors of other users, To control the temperature parameter of learning difficulty, The cosine similarity function is used to measure the similarity between two feature vectors; S34. Use the constructed loss function to train the model. During the training process, the original feature vectors and augmented feature vectors of users in the unlabeled data set are input into the model, the loss is calculated, and the model parameters are updated through back propagation to optimize the model weights, so that the model can learn common transaction patterns and behavior rules from the unlabeled data; S35. After the training is completed, a small amount of annotated abnormal transaction data is introduced to fine-tune the pre-trained model, and the annotated data is input into the model. The model parameters are further optimized through the cross loss function to adapt the model to the specific needs of the abnormal transaction scenario of the game, and the abnormal probability score is output. The formula is: , where is the total cross entropy loss of the entire annotation technique, is the number of samples in the annotated abnormal transaction data, For the The true labels of samples, For the model Samples The output anomaly probability score, For the The transaction data characteristics of the input; S36. Abnormal probability score based on model output , set the threshold ,when When the user's transaction behavior is judged to be abnormal, At this time, the user's transaction behavior is judged to be normal.

5. A method for identifying abnormal transaction behavior of game users according to claim 4, characterized in that: In step S32, generating an augmented feature vector includes the following steps: S321, randomly replace part of the information in the transaction item type with special tags, forcing the model to learn to recover complete features from incomplete information; S322. Add random time offsets to transaction timestamps to simulate small shifts in the time series and enhance the robustness of the model to time changes. S323. Randomly sample some subgraphs from the social relationship graph of the reporting user, retain key social links, remove honor information, and make the model focus on core social features.

6. A method for identifying abnormal transaction behavior of game users according to claim 4, characterized in that: In step S4, processing the transaction data and classifying the risk levels includes the following steps: S41, using a stream processing framework to access game transaction records in real time, including but not limited to the timestamp of each transaction, the IDs of both transaction parties, the transaction amount, and the transaction item information, to generate a real-time feature vector; S42, using edge computing nodes to pre-process transaction data locally in the server cluster, extract lightweight graph features related to the device environment, including but not limited to device fingerprints, IP addresses, and whether a simulator is used, and generate a unique identifier; S43, compare the real-time feature vector with the dynamic behavior baseline and calculate the deviation. The formula is: , where is the real-time feature vector No. eigenvalues, is the behavioral baseline vector No. eigenvalues, is the feature dimension, when The larger the value of, the higher the degree to which the current transaction behavior deviates from the normal mode; S44. Integrate the real-time transaction data into the dynamic transaction relationship graph, update the graph structure and node features, analyze the relationship graph, and output the abnormal score of each transaction behavior. The formula is: , where For Node The anomaly score, is the dimension of the embedding vector, and Node The embedding vector of the abnormal transaction cluster and the centroid vector of the Quantity; S45. Calculate the comprehensive risk score based on the anomaly score, combined with the dynamic behavior baseline deviation and the transaction anomaly score. The formula is: , where and as well as are the corresponding weight coefficients, , and divide the risk threshold, and divide the risk level according to the risk score. When When When When , it is classified as extremely high risk, where and as well as The risk thresholds are set as low, medium and high respectively.

7. A method for identifying abnormal transaction behavior of game users according to claim 4, characterized in that: In step S5, generating an abnormal transaction evidence chain and reversely optimizing model parameters and risk scores includes the following steps: S51. Calculate the SHAP value to quantify the contribution of each feature to the model prediction results. The formula is: , where Features SHAP value of is a feature set A subset of For subset The number of elements of is the total number of features, When only a subset is used The features in are the predicted values ​​of the model, To feature Add Subset The model prediction value changes after S52, filtering out features with larger absolute values ​​of SHAP values, and converting key features and their SHAP values ​​into natural language in combination with a preset explanation template; S53. Evidence data is collected and hashed to generate a unique transaction hash value. The formula is: , where is the transaction hash value, For data splicing, , , , , as well as The transaction timestamp, transaction amount, transaction item information, device fingerprint, SHAP value interpretation label and output abnormal probability score are recorded in the blockchain distributed ledger to ensure the authenticity and permanent storage of the evidence chain. S54. The system displays the preliminary determination results of abnormal transactions, SHAP value interpretation labels, and original evidence data to the manual reviewer for manual determination and annotation, and recycles the manually annotated data; S55, the recovered labeled data is divided into a training set and a test set, and the gradient descent optimization algorithm is used to back-propagate the model parameters. The formula is: , where is the eigenvector value of the cross loss function after back propagation, For sample The real label, For the model to sample The predicted probability of .

8. A method for identifying abnormal transaction behavior of game users according to claim 1, characterized in that: In step S6, optimizing the strategic network parameters includes the following steps: S61. Integrate the current game's economic status, user behavior data, and model recognition results into a state vector, and the formula is: , where is the integrated state vector, is a vector of game economic indicators, including but not limited to the inflation rate of game currency and the price deviation of rare equipment. is the deviation of user behavior baseline, is the historical risk level sequence, is the abnormal probability score output by the model; S62. Define executable detection strategy actions and design a reward mechanism based on abnormal transaction determination and optimization results. The formula is: , where is the reward value, is the rate of change of abnormal transaction identification accuracy, It is the improvement value of the stability of the game economic indicators. is the number of false positives, , and are the corresponding weights respectively; S63, constructing a policy network, outputting the probability distribution of each action, and optimizing the policy network parameters according to the dynamic strategy; S64. Repeat steps S61-S63, continuously interact with the game environment, accumulate experience data and update the strategy network until the loss function converges or the reward value reaches a stable state, and then output the final dynamic detection strategy.

9. A method for identifying abnormal transaction behavior of game users according to claim 8, characterized in that: In step S63, optimizing the policy network parameters includes the following steps: S631. Use a deep neural network to build a policy network, input a state vector, and output the probability distribution of the action. The formula is: , where Strategy Network In a given state Take action The probability of is the parameter set of the policy network, and is the network weight matrix, and is the bias vector; S632, randomly initialize the policy network parameters, and set the learning rate and discount factor parameters; S633, at each time step, according to the current state, sampling actions through the strategy network, and performing corresponding detection strategy adjustments; S634: After executing the sampling action, observe the changes in the game environment and obtain a new state and reward value; S635. Use the temporal difference learning algorithm to define the objective function, the formula is: , where For the strategy Next, status Execute action when The state-action value function, To find the expectation of the internal random variable, For The reward value obtained at any time, is the discount factor, with a value range of [0-1]. The closer its value is to 0, the more emphasis is placed on current rewards, and the closer its value is to 1, the more emphasis is placed on future rewards. For the strategy Next, next state Execute the next action The maximum state-action value of S636. Optimize the network strategy by minimizing the loss function. The formula is: , where is the loss function value, Strategy Network In a given state Take action Logarithm of probability; S637, use stochastic gradient descent to update the network strategy, the formula is: , where is the learning rate, is the loss function Parameters gradient.

10. A system for identifying abnormal transaction behavior of game users, applied to a method for identifying abnormal transaction behavior of game users according to any one of claims 1 to 9, characterized in that: include: Data acquisition and processing module, used to collect multi-source data and pre-process the data; Dynamic transaction building module, used to build a dynamic transaction model for users based on processed data; Self-supervised learning module, used to mine the structure and rules within the data without manual annotation; The risk assessment and classification module uses the output feature representation to conduct risk assessment on user transaction behaviors; The traceability layer processing module is used to conduct traceability analysis on abnormal transactions that are judged to be high-risk; The reinforcement learning strategy module is used to continuously optimize the detection strategy based on system feedback and dynamically adjust model parameters and detection rules.

Citation Information

Patent Citations

  • Multi-agent, multi-user and multi-terminal collaborative smart park cloud service system

    CN110336886A

  • Intelligent park operation index calculation system based on digital twinning

    CN117521969A

  • Systems and methods for risk factor predictive modeling with model explanations

    US11983777B1

  • Predictive Model Data Stream Prioritization

    US20230123322A1

Cited By

  • Game behavior anomaly detection method based on spatio-temporal feature fusion

    CN120705712A

  • A method for detecting game behavior anomalies based on spatiotemporal feature fusion

    CN120705712B

  • Campus second-hand commodity transaction credible lake source evaluation method and system

    CN121352802A