AI-based gift bag substitution identification and risk control method, system and equipment
By constructing an AI dynamic risk control model and combining multi-dimensional data and gift pack lifecycle analysis, the shortcomings of traditional risk control methods in dealing with malicious behavior of black market organizations have been solved. This has enabled accurate identification and real-time interception of gift pack top-ups, improving the safe operation and transaction security of game item gift packs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI SANQI JIYU NETWORK TECH CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional risk control methods are insufficient to effectively address the complex and ever-changing malicious behaviors of black market organizations, threatening the security of game item gift pack operations and the interests of the platform. Existing technologies lack flexibility and adaptability, making it difficult to achieve accurate identification and real-time interception.
We construct an AI dynamic risk control model based on multi-dimensional data and gift package lifecycle analysis. By building a feature warehouse, configuring combined risk rules, integrating trust scoring and anomaly detection models, and using deep Q-networks for real-time adjustments and SHAP models for attribution analysis, we achieve accurate identification and interception.
It enables accurate identification and real-time interception of gift pack top-up behavior, improves the safe operation of game item gift packs, ensures transaction security, and enhances the scientific nature and explainability of risk control decisions.
Smart Images

Figure CN121998702A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an AI-based method, system, and device for identifying and controlling the risk of gift package top-ups. Background Technology
[0002] In the long-term operation of a game, item packs, as core assets and an important vehicle for user payments, play a crucial role in establishing trust between the platform and users. They are not only an important component of the game's economic system, directly affecting its balance and enjoyment, but also a vital means for game operators to generate revenue and sustain the game's continued development.
[0003] However, the current operation of game item gift packs faces severe challenges, with malicious activities by black market organizations seriously damaging the game ecosystem and platform interests. These organizations employ various methods for illegal operations, such as maliciously exploiting event vulnerabilities to obtain large amounts of gift packs that should be distributed to users, thus undermining the fairness of the events; using code vulnerabilities to bypass game rules and restrictions for illegitimate gift pack acquisition and trading; and offering recharge services at low prices, such as obtaining a gift pack worth 648 yuan for only 100 yuan, profiting from the difference, severely disrupting the game's economic order.
[0004] Specifically, criminal organizations also generate a large number of fake accounts in a short period of time by registering accounts in bulk, increasing instability within the game; they forge IP addresses to circumvent IP address-based risk control detection, making it difficult for the risk control system to accurately identify their real location and behavior patterns; and they engage in cross-regional top-ups, taking advantage of the differences between game servers in different regions to conduct illegal top-up operations and obtain illicit profits. These methods are complex and varied, causing great trouble for game operations.
[0005] In the face of these malicious acts by black market organizations, traditional risk control methods mainly rely on a single rule engine or manual review. However, these methods have many limitations and are difficult to effectively deal with the complex and ever-changing black market tactics.
[0006] Single-rule engine-based risk control methods typically rely on rules based on known patterns of illicit activity, lacking flexibility and adaptability. When illicit organizations change their behavior or adopt new methods, existing rules may fail to identify and block them in a timely manner, significantly reducing the effectiveness of risk control. Moreover, single-rule engines often only make judgments from a limited perspective, making it difficult to fully consider the complex relationships between various factors, which can easily lead to misjudgments or omissions.
[0007] While manual review can identify suspicious behavior to some extent based on the experience and judgment of reviewers, it is inefficient and costly. With the continuous increase in the number of game users and the ever-growing volume of transactions, manual review cannot meet the needs of real-time risk control and struggles to conduct timely and accurate reviews of every transaction. Furthermore, manual review is highly subjective; different reviewers may have different judgments on the same behavior, leading to inconsistent review results.
[0008] Therefore, traditional risk control methods are no longer sufficient to deal with the malicious behavior of black market organizations and cannot effectively protect the safe operation of game item gift packs and the interests of the platform. Summary of the Invention
[0009] The purpose of this invention is to provide an AI-based method, system, and device for identifying and controlling the risk of gift pack top-ups. By constructing an AI dynamic risk control model based on multi-dimensional data and gift pack lifecycle analysis, it can effectively solve the problems existing in traditional risk control methods, achieve accurate identification and real-time interception of gift pack top-up behavior, and provide strong protection for the safe operation of game item gift packs, thereby solving at least one of the aforementioned problems in the prior art.
[0010] In a first aspect, the present invention provides an AI-based method for identifying and controlling the risk of third-party top-up services using gift packages, the method specifically comprising: A feature repository is constructed based on the lifecycle data of gift packs and marketing data from game servers, combined with account, device, and network behavior data collected by security components. Based on the feature warehouse, a set of combined risk rules containing multi-dimensional factors are configured. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, a primary risk label is triggered. The initial risk markers are input into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the initial risk markers, a comprehensive risk score and handling recommendations are output. The results of comprehensive risk scoring, implementation and disposal recommendations, and their feedback are used as state inputs to a deep Q-network, which automatically adjusts the risk scoring threshold, disposal action intensity, and weights of combined risk rules based on real-time risk control effectiveness indicators. The SHAP model interpretability technique is used to perform attribution analysis on the model decision-making process of high-risk cases and output the contribution of key risk features.
[0011] Secondly, the present invention provides an AI-based system for identifying and controlling the risk of gift package top-ups, the system specifically comprising: The data acquisition module is used to build a feature repository based on the lifecycle data of gift packs and marketing data from the game server, combined with account, device and network behavior data collected by the security component. The rule configuration module is used to configure a set of combined risk rules containing multi-dimensional factors based on the feature warehouse. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, a primary risk label is triggered. The risk scoring module is used to input the primary risk markers into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the primary risk markers, it outputs a comprehensive risk score and handling recommendations. The risk adjustment module is used to input the results of the comprehensive risk score, the implementation and disposal suggestions and their feedback into the deep Q network, and automatically adjust the risk score threshold, the intensity of disposal actions and the weight of combined risk rules according to the real-time risk control effect indicators. The attribution analysis module is used to perform attribution analysis on the model decision-making process of high-risk cases using SHAP model interpretability technology, and output the contribution of key risk features.
[0012] Thirdly, the present invention provides a computer device, including: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the AI-based gift package top-up identification and risk control method as described in any of the above methods.
[0013] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the AI-based gift package top-up identification and risk control method as described in any of the above methods.
[0014] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention constructs an AI dynamic risk control model based on multi-dimensional data and gift pack lifecycle analysis, which can effectively solve the problems existing in traditional risk control methods, realize the accurate identification and real-time interception of gift pack recharge behavior, and provide strong protection for the safe operation of game item gift packs.
[0015] 2. This invention integrates multi-source data to build a feature warehouse and combines multiple model outputs to achieve accurate identification and risk control of gift pack top-up services, effectively ensuring the security of game transactions.
[0016] 3. This invention is based on the configuration of combined risk rules with quantitative features and real-time matching and judgment, which can quickly and accurately trigger primary risk marking and capture potential risks in a timely manner.
[0017] 4. This invention integrates the outputs of multiple models to form a comprehensive risk score and disposal suggestions, comprehensively assesses risks and provides reasonable response strategies, thereby improving the scientific nature of risk control decisions.
[0018] 5. This invention constructs a multi-task learning model and combines it with real-time fine-tuning, which can dynamically output accurate user trust values and accurately reflect the stability and value of users' historical behavior.
[0019] 6. This invention trains a dynamic baseline model and calculates the degree of deviation to output anomaly scores, which can effectively identify anomalies in the current transaction chain and accurately assess transaction risks.
[0020] 7. This invention inputs multiple types of data into a deep Q-network and adjusts parameters based on real-time indicators to achieve dynamic adaptive optimization of risk control parameters and improve risk control effectiveness.
[0021] 8. This invention uses SHAP technology to perform attribution analysis on high-risk cases and generate multimodal reports, enhancing the interpretability of model decisions and facilitating understanding and decision-making. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating an AI-based method for identifying and controlling the risk of gift pack top-ups, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of an AI-based gift package top-up identification and risk control system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0028] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating an AI-based method for identifying and controlling the risk of gift package top-ups, as disclosed in an embodiment of the present invention, is shown below in detail: S101 builds a feature repository based on the lifecycle data of gift packs and marketing data from the game server, combined with account, device, and network behavior data collected by security components.
[0031] In this embodiment, the game server log system collects real-time data on the entire lifecycle of the gift pack from creation to expiration, including the gift pack type (such as first recharge gift pack, holiday limited gift pack), generation time, distribution channel (such as event page, mail system), claim conditions (such as level restriction, task completion status), claim time, usage status (unused / used / expired) and post-use effect data (such as changes in item attributes, increases or decreases in virtual currency).
[0032] Retrieve marketing campaign data associated with gift packs from the game marketing backend system, including campaign type (e.g., limited-time discount, spending threshold), campaign time range, participating user groups (e.g., new users / existing users), campaign rules (e.g., discount percentage, spending threshold), and campaign performance data (e.g., gift pack sales, user participation rate). Perform correlation analysis on the marketing data, mapping gift pack IDs to campaign IDs to form a "gift pack-campaign" association table, and label the role of each gift pack in the marketing campaign (e.g., main gift pack, supplementary gift pack).
[0033] User account behavior data is collected through security components (such as the game's built-in anti-cheat SDK), including account registration information (registration time, registration channel, bound device ID), login behavior (login time, login frequency, login IP address, login device model), recharge behavior (recharge amount, recharge channel, recharge time), and social behavior (number of friends, guild membership status, chat history keywords). Account data is anonymized, hiding sensitive user information (such as real name and mobile phone number), retaining only identifiers related to behavioral characteristics.
[0034] Security components are used to collect user device and network environment data, including device hardware information (device model, operating system version, IMEI / MAC address), device environmental data (screen resolution, battery status, sensor data), network connection information (IP address, network type, carrier, network latency), and geographic location data (rough geographic location based on IP positioning, precise geographic location based on GPS positioning). Device data is standardized by mapping device models to unified category labels (e.g., "high-end / mid-range / low-end") and classifying network latency into tiers (e.g., "low latency / medium latency / high latency").
[0035] Using user account IDs as unique identifiers, the system integrates and correlates gift pack lifecycle data, marketing data, account behavior data, and device and network behavior data. For example, it aligns the time a user claims a gift pack with their login and recharge times on a timeline to analyze user behavior patterns before claiming the gift pack; and it correlates the effectiveness of gift pack usage with device performance data to analyze the consumption preferences of users on different devices for gift pack items.
[0036] Based on user behavior time series, user session data is constructed. A session is defined as a complete interaction process from logging into the game to logging out, recording all behavioral events within the session (such as claiming gift packs, using items, participating in activities, and making in-game purchases) and the temporal relationships between these events. The session data is segmented, for example, into 5-minute windows, to analyze the potential risks of intensive behaviors within a short period (such as claiming gift packs in bulk or making rapid in-game purchases).
[0037] Higher-order combined features are generated through feature cross-referencing. For example, cross-referencing "account registration channel" with "gift pack redemption channel" generates the "registration channel - redemption channel" combined feature, which analyzes the preferences of users with different registration channels for gift pack redemption channels; cross-referencing "device performance tiers" with "gift pack usage effect" generates the "device performance - item consumption speed" combined feature, which identifies the correlation between device performance and gift pack usage behavior.
[0038] Basic features are extracted from the merged data, including statistical features (such as the total number of gift packs a user has received in history and the highest amount recharged in a single day), time-series features (such as the login frequency in the last 7 days and the distribution of gift pack receiving time), distribution features (such as the frequency distribution of recharge amount and the skewness of the distribution of gift pack usage time), and relationship features (such as the proportion of friends who received the same gift pack and the average recharge amount of guild members).
[0039] Based on the basic features, further derived features are generated. For example, recharge fluctuation features are generated by calculating the "deviation rate between daily recharge amount and historical average," claim timeliness features are generated by calculating the "interval between gift pack claim time and event start time," and geographical location consistency features are generated by comparing "device geographical location and account registration geographical location." The derived features are then normalized, with numerical features scaled to the [0,1] interval, and categorical features converted to one-hot encoding.
[0040] Features are stored in layers based on business scenarios and update frequency. The basic feature layer stores static user attributes (such as registration information and device information) and low-frequency update features (such as historical cumulative recharge amount), updated daily. The dynamic feature layer stores high-frequency user behavior features (such as login count in the last hour and gift pack redemption records in the last 10 minutes), updated in real time. The tag feature layer stores risk tags (such as whether the account has been marked as a black market account or whether it has participated in abnormal activities), updated dynamically based on risk control decision results. This layered storage enables efficient feature querying and version management.
[0041] Regularly evaluate the quality of features in the feature repository, including feature coverage (the proportion of users / behaviors covered), feature stability (the degree of fluctuation in feature values over time), feature discriminative power (the significance of feature value differences between positive and negative samples), and feature redundancy (the correlation coefficient between features). Eliminate or optimize low-quality features, such as merging highly correlated features, supplementing missing value imputation strategies, and adjusting feature calculation logic, to ensure the conciseness and effectiveness of the feature repository.
[0042] This embodiment constructs a feature repository containing multi-dimensional and multi-level features, providing data support for subsequent configuration of combined risk rules and model training. This feature repository supports dynamic expansion, continuously incorporating newly collected data and newly generated features as the gaming business develops and black market methods evolve, maintaining its ability to identify risky behaviors.
[0043] S102, based on the feature warehouse, configure a set of combined risk rules containing multi-dimensional factors. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, the primary risk label is triggered.
[0044] In this embodiment, key dimensions are extracted and categorized from a feature repository based on the typical characteristics of black market activities. For example, the account dimension may include: the proportion of new accounts (e.g., the proportion of accounts registered within 24 hours), abnormal login behavior (e.g., cross-regional login within a short period), and account transaction frequency (e.g., the number of times recharges or gift packs are claimed within a unit of time); the device dimension may include: device fingerprint duplication rate (e.g., multiple accounts registered on the same device), and abnormal device environment (e.g., emulator or virtual machine environment); the network dimension may include: IP address clustering (e.g., a large number of accounts operating under the same IP), and abnormal network request frequency (e.g., exceeding the normal user request rate); the gift pack transaction dimension may include: abnormal transaction amount (e.g., low-price recharge, excessive discounts), concentrated transaction time (e.g., high-frequency transactions during inactive periods), and abnormal gift pack usage path (e.g., obtaining gift packs without participating in activities). Specific indicators are further refined under each dimension to form a quantifiable risk factor library.
[0045] Based on multi-dimensional factors, a logical framework for combined risk rules is designed. The rules must meet two core conditions: "multi-factor collaborative judgment" and "time window restriction". For example, for the top-up behavior, the following rules can be configured: if an account meets the following conditions within 10 minutes (a specific time window), a primary risk mark will be triggered: (1) Account dimension: registration time is less than 24 hours and real-name authentication has not been completed; (2) Device dimension: device fingerprint matches the historical black market device database; (3) Network dimension: the IP address location is inconsistent with the account registration location and there are more than 5 active accounts under the IP address at the same time; (4) Transaction dimension: the single top-up amount is 648 yuan (the highest tier in the game) but the actual payment amount is less than 100 yuan. When configuring the rules, the weight threshold of each dimension (such as the abnormality level of a certain dimension must reach the "high risk" level) and the logical relationship between dimensions (such as "AND" relationship or "OR" relationship) should be clearly defined to ensure that the rules can accurately capture complex black market behaviors.
[0046] The setting of time windows needs to be dynamically adjusted based on the temporal characteristics of black market activities and normal user behavior patterns. For example, for bulk account registration, a time window condition can be set as "more than 3 accounts registered on the same device within 5 minutes"; for malicious exploitation of promotional benefits, a time window condition can be set as "the number of times a single account claims a gift pack within the first 30 minutes after the event starts, exceeding twice the maximum limit set by the event rules". Simultaneously, by analyzing historical data on peak periods of black market activity (such as 2-5 AM), the time window threshold for these periods can be tightened (e.g., shortening the time span or lowering the dimensional anomaly threshold), while the rules can be appropriately relaxed during normal user activity periods to reduce false positives. Optimizing time windows requires continuous monitoring of rule triggering frequency and the effectiveness of black market interception, and parameter adjustments can be made through A / B testing.
[0047] When multiple factors simultaneously meet the abnormal conditions of the combined risk rules within a specific time window, the system automatically triggers a primary risk marker. The marker must include: the trigger rule number, risk type (e.g., third-party top-up, fraudulent activity, cross-regional transactions), involved account ID, device fingerprint, IP address, timestamp, and details of abnormal indicators for each dimension (e.g., registration time, transaction amount, IP concentration, etc.). After the primary risk marker is generated, it must be immediately recorded in the risk event database and simultaneously pushed to subsequent risk control modules (e.g., trust scoring models) for further analysis. Simultaneously, to avoid duplicate markers, multiple triggers by the same account within a short period (e.g., within 1 hour) must be deduplicated, retaining only the marker with the highest risk level.
[0048] S103. Input the primary risk marker into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the primary risk marker, output a comprehensive risk score and disposal recommendations.
[0049] In this embodiment, after the system triggers a primary risk marker, the marked data needs to be preprocessed to ensure the standardization of the model input. First, the core fields in the marker are extracted, including account ID, risk type (e.g., third-party top-up, fraudulent activity, cross-regional activity), trigger rule number, timestamp, and details of abnormal indicators for each dimension (e.g., registration time, transaction amount, IP clustering). Second, unstructured data (e.g., device fingerprints, IP addresses) is standardized; for example, device fingerprints are mapped to unique codes, and IP addresses are converted to geographic coordinates and network service provider information. Finally, continuous numerical indicators (e.g., transaction amount, operation frequency) are normalized to map to the [0,1] interval, eliminating the impact of dimensional differences on the model. The preprocessed data is stored in a structured format and simultaneously pushed to the input interfaces of the trust scoring model and the anomaly detection model.
[0050] The trust scoring model dynamically assesses an account's trust level based on its historical behavior data and real-time risk markers. The model first extracts long-term behavioral features from a feature repository, including registration duration, number of historical transactions, frequency of normal activity participation, and number of historical risk markers, to construct a static trust baseline. For example, accounts registered for more than 30 days and with no historical risk records can be assigned a higher initial trust value. Subsequently, the static trust baseline is dynamically adjusted based on the current initial risk marker rating (e.g., low, medium, high): if an account triggers a high-risk marker (e.g., third-party top-up behavior), its trust value is reduced according to preset rules (e.g., 20% trust value deduction for each high-risk marker); if the account does not trigger risks again in subsequent activities, its trust value is gradually restored according to a time decay factor (e.g., 5% trust value restoration every 24 hours). The final output dynamic trust value is a value within the range [0, 100], with higher values indicating higher account credibility.
[0051] Anomaly detection models identify the degree of anomalousness in account behavior using unsupervised learning algorithms (such as Isolation Forest or cluster analysis). Model input includes real-time behavioral characteristics of the account (such as current transaction amount, operation frequency, and device environment) and anomaly indicators from the primary risk markers (such as IP clustering and device fingerprint duplication rate). First, the model compares the account's real-time behavioral characteristics with historical normal behavior patterns, calculating the behavioral deviation (e.g., whether the transaction amount exceeds three standard deviations of the historical mean). Second, it combines the anomaly indicators from the primary risk markers (e.g., the presence of more than five active accounts under the same IP) to weight the deviation (e.g., anomaly in IP clustering can increase the deviation weight by 20%). Finally, the model outputs an anomaly score, ranging from [0,1], with scores closer to 1 indicating a higher degree of behavioral anomalousness. For example, an account with an anomaly score of 0.9 may be involved in bulk registration or third-party top-up activities.
[0052] The comprehensive risk score is generated by integrating dynamic trust value, anomaly score, and primary risk marker rating. First, weights are assigned to each dimension: dynamic trust value is weighted at 40% (reflecting the account's historical credibility), anomaly score at 40% (reflecting the degree of current behavioral abnormality), and primary risk marker rating at 20% (reflecting the direct risk signal from rule matching). For example, if an account has a dynamic trust value of 70 (medium credibility), an anomaly score of 0.8 (high anomaly), and a primary risk marker rating of "high," its comprehensive risk score is calculated as: 70 × 40% + 0.8 × 100 × 40% + 100 (score corresponding to high risk level) × 20% = 28 + 32 + 20 = 80. The comprehensive risk score ranges from [0, 100], with higher scores indicating higher risk. Accounts with scores exceeding 60 are marked as high-risk and require priority handling.
[0053] Based on the comprehensive risk score, the system generates differentiated handling suggestions. For low-risk accounts (score < 30), an "observation" strategy is adopted, recording risk events but not restricting account functions; for medium-risk accounts (30 ≤ score < 60), a "transaction restriction" strategy is adopted, such as temporarily prohibiting the collection of gift packs or recharges, and pushing for manual review; for high-risk accounts (score ≥ 60), a "forced intervention" strategy is adopted, including freezing the account, rolling back abnormal transactions, and adding to the blacklist. Simultaneously, the system further refines the handling suggestions according to the risk type: for example, for recharge on behalf of others, the flow of funds needs to be traced and associated accounts banned; for cross-region recharges, server permissions need to be adjusted and affected users compensated. After the handling suggestions are generated, they are sorted from high to low risk scores, prioritizing the handling of high-risk events to ensure efficient allocation of risk control resources.
[0054] After implementing the suggested actions, the system records the results (such as whether the account is unblocked, user appeal status, and the effectiveness of black market interception) and feeds the results back to the trust scoring model and anomaly detection model. For example, if an account does not exhibit any risky behavior after being frozen, its dynamic trust value can be gradually restored; if a rule-triggered action suggestion leads to a large number of misjudgments, the weight or threshold of that rule needs to be adjusted. In addition, the system regularly analyzes the handling data of high-risk cases, extracts new risk features (such as device fingerprint patterns of new types of top-up methods), updates the feature repository and combined risk rules, forming a closed-loop iterative mechanism of "detection-handling-optimization" to continuously improve the accuracy and adaptability of the risk control model.
[0055] S104 uses the results of comprehensive risk scoring, implementation and disposal recommendations, and their feedback as state inputs to a deep Q-network, which automatically adjusts the risk scoring threshold, disposal action intensity, and weights of combined risk rules based on real-time risk control effectiveness indicators.
[0056] In this embodiment, the comprehensive risk score, the execution results of the handling suggestions, and their feedback data are integrated into a unified state input. First, the comprehensive risk score, as the core indicator, is divided into discrete states according to a preset range (e.g., low risk [0-30], medium risk [31-60], high risk [61-100]) to simplify subsequent network input. Second, the execution results of the handling suggestions include the handling type (e.g., observation, transaction restriction, account freezing), execution timestamp, and post-handling account behavior data (e.g., whether the risk rule is triggered again). These data are converted into structured fields through feature engineering; for example, the handling type is encoded as a numerical value (observation=1, transaction restriction=2, account freezing=3). Finally, the feedback data covers indicators such as user appeal status, black market interception success rate, and false positive rate, which are normalized and mapped to the [0,1] range to eliminate dimensional differences. The integrated state data is stored in time series form, forming continuous input samples to ensure that the deep Q network can capture the dynamic changes of risk control strategies.
[0057] The system defines a set of real-time risk control effectiveness metrics to evaluate the effectiveness of the current strategy. These metrics include interception rate, false positive rate, user satisfaction, and black market activity.
[0058] Deep Q-networks take state data as input and output the Q-value (expected long-term reward) of each selectable action. The network structure uses a multilayer perceptron (MLP), with the number of input layer nodes matching the dimensions of the state data (e.g., 10 features such as comprehensive risk score, treatment type, and interception rate), 3-5 hidden layers to capture non-linear relationships, and the number of output layer nodes corresponding to the number of selectable actions (e.g., adjusting the risk score threshold, increasing / decreasing the intensity of treatment actions, and modifying the weights of combination rules).
[0059] During training, the system samples state-action-reward triplets from historical risk control logs (e.g., in a certain state, selecting "increase the risk scoring threshold" increases the interception rate but also the false positive rate, with a combined reward of +0.2). Through an experience replay mechanism, it breaks down data correlations and improves training stability. The network updates its parameters by minimizing a loss function (e.g., mean squared error), gradually bringing the Q-value closer to the true reward, ultimately forming a mapping from state to optimal action. For example, when the state shows "low interception rate but high false positive rate for high-risk accounts," the network might output that "lowering the risk scoring threshold and reducing the intensity of the freeze action" yields the highest Q-value, serving as a basis for policy adjustment.
[0060] Based on the action suggestions output by the Deep Q network, the risk score threshold is dynamically adjusted. The initial threshold is set as an empirical value (e.g., a comprehensive risk score ≥ 60 is marked as high risk). When the network suggests adjusting the threshold based on real-time status analysis (e.g., the current false positive rate is too high), the system modifies the threshold by a preset step size (e.g., ±5 points). For example, if the status shows "the proportion of medium-risk accounts mistakenly restricted from trading exceeds the threshold," the network may suggest lowering the high-risk threshold from 60 to 55 to reduce false positives; conversely, if black market activity increases, the network may suggest raising the threshold from 60 to 65 to improve interception accuracy. The adjusted threshold takes effect immediately and is fed back to the network as part of the new status, forming a closed-loop optimization.
[0061] The intensity of the action (such as the duration of transaction restrictions and the severity of account freezes) is dynamically adjusted based on network suggestions. The system predefines action intensity levels (e.g., transaction restrictions are divided into three levels: 1 hour, 24 hours, and 72 hours), and the network outputs intensity adjustment suggestions based on the status. For example, when the status shows "High-risk accounts have a high rate of repeated violations," the network may suggest upgrading "transaction restriction" to "account freeze"; if the status shows "User appeals are concentrated on freezing actions," the network may suggest downgrading "account freeze" to "24-hour transaction restriction." The adjusted action intensity is directly applied to subsequent processing procedures, and the effect is verified through feedback data (such as changes in the appeal rate) to continuously optimize the action selection strategy.
[0062] The combined risk rules are composed of weighted factors from multiple dimensions (such as device anomaly, IP clustering, and transaction amount deviation). The network suggests adjusting the weights of each dimension based on the status. The initial weights are set based on historical black market patterns (e.g., device anomaly weight 30%, IP clustering weight 20%). When the network suggests adjusting the weights based on real-time status analysis (e.g., a certain type of black market activity shifting to mass IP spoofing), the system reallocates the weights proportionally. For example, if the status shows "IP spoofing leads to an increase in risk control bypass rate," the network may suggest increasing the IP clustering weight from 20% to 35% while decreasing the device anomaly weight to 15% to strengthen the identification of new types of attacks. After the weights are updated, the system recalculates the account's overall risk score to ensure rule adaptability.
[0063] The effectiveness of dynamically adjusted risk control is regularly evaluated (e.g., daily / weekly calculation of interception rate, false positive rate, etc.). If the effect does not meet expectations (e.g., a decrease in interception rate or an increase in false positive rate), the network is retrained. Training data includes the latest state-action-reward samples to ensure the network learns the latest black market patterns. For example, if black market operators begin to use a "low-frequency, high-amount" top-up strategy, causing existing rules to miss detections, the system adds an "abnormal transaction frequency" feature and retrains the network, prompting it to output a suggestion to "increase the weight of transaction amount deviation," ultimately optimizing the combined rules. Through continuous iteration, the system gradually forms a dynamic risk control strategy that adapts to the evolution of black market activities, ensuring the security of gift package operations.
[0064] S105 uses the SHAP model interpretability technique to perform attribution analysis on the model decision-making process of high-risk cases and outputs the contribution of key risk characteristics.
[0065] In this embodiment, cases marked as high-risk are screened from the real-time risk control process. The screening criteria are a comprehensive risk score exceeding a preset threshold (e.g., ≥85 points) or triggering a high-level handling recommendation (e.g., account freezing). For each high-risk case, the system extracts its associated raw feature data from the feature repository, including account behavior features (e.g., login frequency, transaction amount fluctuations), device features (e.g., device fingerprint uniqueness, hardware information anomalies), network features (e.g., IP address clustering, cross-region recharge records), and marketing activity participation features (e.g., gift pack redemption frequency, number of activity rule triggers). All feature data must be cleaned and standardized, for example, discrete features (e.g., device type) are converted into one-hot encoding, and continuous features (e.g., transaction amount) are normalized to the [0,1] interval to ensure a unified data format and the accuracy of interpretable analysis.
[0066] Initialize the SHAP (SHapley Additive exPlanations) model, which is based on the Shapley value theory in game theory. It quantifies the importance of features by calculating the marginal contribution of each feature to the model output. During initialization, a baseline value must be defined (i.e., the model's prediction result when all features are averaged or at their default values) as a reference point for subsequent contribution calculations. For example, when analyzing the overall risk score of a high-risk account, the baseline value can be set as the average risk score of all normal accounts (e.g., 40 points). By comparing the actual score (e.g., 90 points) with the baseline value, the contribution of each feature can be decomposed.
[0067] For each high-risk case, the SHAP model calculates the feature contribution through the following steps: (1) Sampling and perturbation: Randomly sample some feature combinations from the feature warehouse (such as randomly hiding some device features or network features) to generate multiple perturbation samples to simulate the changes in model output when features are missing or changed; (2) Marginal contribution calculation: For each perturbation sample, calculate the difference between its model output (such as the comprehensive risk score) and the baseline value as the marginal contribution of the current feature combination; (3) Weighted average: According to the frequency of the feature in all sampled combinations, perform a weighted average of the marginal contribution to obtain the final contribution of each feature. For example, if the "device fingerprint uniqueness" feature leads to a significant increase in risk score in most high-risk cases (such as contribution +20 points), it is marked as a key risk feature.
[0068] The contribution results of all high-risk cases are aggregated, and the average contribution and frequency of each feature in the overall cases are calculated. The top N features (e.g., the top 5) in terms of contribution are selected as key risk features.
[0069] Attribution analysis is performed on the aggregated key risk features, and their impact paths are explained in conjunction with business logic. For example, if the "cross-region recharge record" feature has the highest contribution (e.g., an average of +25 points), the attribution analysis indicates that black market operators may be using exchange rate differences between different servers to conduct low-price recharges; if the "device fingerprint uniqueness" feature has a prominent contribution (e.g., an average of +18 points), it is attributed to black market operators using virtual devices in bulk to bypass risk control detection. The system generates a visual report, displaying the contribution ranking of each feature in the form of bar charts or heatmaps, and annotating the specific meaning and typical cases next to each feature (e.g., "Account A has a risk score of 95 points because it simultaneously possesses the features of 'cross-region recharge' and 'IP aggregation'"), to help operations personnel quickly understand the model's decision-making logic.
[0070] The attribution analysis results are fed back to the risk control strategy layer to optimize combined risk rules and model parameters. For example, if the "transaction amount fluctuation" feature has a low contribution in multiple high-risk cases (e.g., an average of +3 points), but this feature is often mistakenly amplified in misjudged cases, its weight can be adjusted or related features can be added (e.g., the "transaction amount fluctuation + multiple transactions in a short period of time" combined rule). If the "abnormal device hardware information" feature has a significant contribution (e.g., an average of +22 points), the accuracy of the device fingerprint detection module can be improved. In addition, the attribution results can also be used to train trust scoring models and anomaly detection models. For example, high-contribution features can be used as strong association rules input into the model to improve its ability to identify new black market methods.
[0071] A dynamic monitoring mechanism is established to periodically (e.g., daily / weekly) re-screen high-risk cases and perform attribution analysis to track the changing trends of key risk characteristics. For example, if the black market begins to use new attack methods (such as using AI to generate virtual accounts), causing a decrease in the contribution of "device fingerprint uniqueness" and an increase in the contribution of "account behavior pattern similarity," the system can automatically adjust the input feature set of the SHAP model, incorporating new features (such as account operation time distribution and task completion path similarity) and recalculating the contribution, ensuring that the attribution analysis always aligns with the latest black market patterns. Through continuous iteration, the system forms a closed loop of "detection-attribution-optimization-re-detection," effectively improving the accuracy and interpretability of risk prevention and control for gift pack top-ups.
[0072] In some embodiments, in step S101 above, the construction of a feature warehouse based on the gift pack lifecycle data and marketing data from the game server, combined with account, device, and network behavior data collected by the security component, specifically includes: Through the application programming interface and data bus, it accesses gift pack transaction logs from the game business server, activity data from the marketing platform, and account, device, and network behavior data collected from the client security component, forming multi-source heterogeneous data; Multi-source heterogeneous data is formatted, cleaned, and standardized to form standardized data. Based on standardized data, users, devices, orders, and marketing activities are extracted as core entities, and the binding, purchase, participation, and asset transfer relationships between core entities are established through order and behavior logs; For core entities, statistical calculations are performed within a preset time sliding window to generate temporal behavioral features and aggregation environment features. Standardized data, relationships, temporal behavioral features, and clustering environment features are concatenated and encoded according to core entities to form a feature warehouse.
[0073] In this embodiment, an application programming interface (API) and a data bus are used to connect to and acquire data from different data sources. Specifically, transaction logs for gift packs are accessed from the game's business server. These logs record all transaction information related to gift packs within the game, including key details such as transaction time, accounts of both parties, and the type and quantity of the gift packs. Activity data is accessed from the marketing platform, covering the rules, timeframes, participation conditions, and rewards for various marketing activities conducted by the game. Simultaneously, account, device, and network behavior data are collected from the client's security component. Account data includes account registration information and login records; device data includes device model, operating system version, and unique device identifier; and network behavior data includes network connection method, IP address, and network access frequency. Through these methods, data from different channels and formats are aggregated to form multi-source heterogeneous data.
[0074] Preprocessing is performed on multi-source heterogeneous data. Because these data come from diverse sources and have varying formats, they often contain missing, incorrect, or duplicate data. Therefore, format unification, cleaning, and standardization are necessary. After these processes, standardized data is generated.
[0075] Based on standardized data, users, devices, orders, and marketing campaigns are extracted as core entities. Users are participants in the game, devices are the tools users use to play the game, orders record transactions between users and the game, and marketing campaigns are important factors influencing user transaction behavior. Relationships between these core entities are established through orders and behavior logs. For example, orders can determine which gift packs a user purchased, thus establishing a purchase relationship between the user and the order; behavior logs can reveal user actions during specific marketing campaigns, establishing a participation relationship between the user and the marketing campaign, and also clarify the asset transfer of gift packs within orders, such as transferring them from the game system to the user's account.
[0076] For core entities, statistical calculations are performed within a preset time sliding window. The time sliding window can be set to different lengths according to actual needs, such as hours, days, or weeks. For user entities, time-series behavioral characteristics such as login frequency, number of transactions, and number of marketing activities participated in are statistically analyzed within the time sliding window. For device entities, characteristics such as network connection stability and usage duration are statistically analyzed within the time period. For order entities, characteristics such as transaction amount distribution and transaction time distribution are statistically analyzed. For marketing activity entities, characteristics such as the number of participants and participation rate are statistically analyzed. Simultaneously, clustered environmental characteristics are considered, such as the behavioral characteristics of user groups in the same region or with the same device model within a certain time period. Through these statistical calculations, time-series behavioral characteristics and clustered environmental characteristics that reflect the behavioral patterns and environmental characteristics of core entities within a specific time period are generated.
[0077] Standardized data, relationships, temporal behavioral features, and clustering environment features are concatenated and encoded according to core entities. The purpose of this concatenation and encoding is to integrate these different types of data into a complete data structure for subsequent risk identification and assessment. For example, for each user entity, account information from its standardized data, purchase order information from its relationships, login frequency and transaction count information from its temporal behavioral features, and user behavior information from its clustering environment features are concatenated and encoded to enable computer systems to recognize and process them. This process ultimately forms a feature repository, providing rich data support for subsequent risk identification and assessment.
[0078] In some embodiments, step S102 above, which involves configuring a set of combined risk rules containing multi-dimensional factors based on a feature warehouse, and triggering a primary risk flag when the multi-dimensional factors simultaneously meet abnormal conditions within a specific time window, specifically includes: Based on the quantitative feature factors defined in the feature warehouse, configure multiple sets of combined risk rules; Upon receiving a transaction risk assessment request, the system queries and assembles the feature vectors in the feature repository corresponding to the transaction risk assessment request in real time to form a assessment context. The judgment context is iterated and matched against all combined risk rules. When the sub-conditions of all dimensions of any combined risk rule are met simultaneously within a specific time window, the combined risk rule is judged to be hit. Based on the pre-configured weights of the hit combined risk rules, a preliminary weighted risk score is calculated and summarized. By comparing the preliminary weighted risk score with the preset risk level threshold, the corresponding primary risk label and its level score are triggered.
[0079] In this embodiment, this work is carried out based on the quantitative feature factors defined in the feature repository. The feature repository stores a wealth of data features related to game item gift pack transactions, covering multiple dimensions such as users, devices, orders, and marketing activities. For example, user-level feature factors may include user login frequency, historical transaction count, and account registration duration; device-level feature factors include device model, device usage time, and device network connection stability; order-level feature factors include order transaction amount, transaction time, and gift pack type; and marketing activity-level feature factors include activity participation count and activity reward acquisition. Based on these quantitative feature factors, combined with common malicious behavior patterns of black market organizations and the actual needs of game operation, multiple sets of combined risk rules are configured. Each set of combined risk rules contains multiple dimensional factors, and corresponding sub-conditions are set for each dimensional factor. For example, a set of combined risk rules targeting recharge monetization behavior may include multiple dimensional factors and their sub-conditions, such as a large number of abnormal recharges in a short period of time in the user dimension, the device being an unused device in the device dimension, and a serious imbalance between the recharge amount and the gift pack value in the order dimension.
[0080] When a new transaction occurs in the game system or a risk assessment of a transaction is required, a transaction risk determination request is issued. Upon receiving the request, the system queries the feature repository in real time to retrieve feature vectors related to the transaction risk determination request. These feature vectors contain detailed feature information related to the user, device, order, etc., related to the current transaction. For example, if the determination is for a user's gift pack recharge transaction, the retrieved feature vectors might include the user's account registration information, historical recharge records, the device information used for the current recharge, the amount recharged, and the gift pack type. These retrieved feature vectors are then assembled according to certain logic to form a complete determination context for subsequent matching with combined risk rules.
[0081] The assembled judgment context is iterated and matched against all pre-configured combined risk rules one by one. During the matching process, for each set of combined risk rules, it is checked whether the sub-conditions of all the dimensions included are simultaneously met within a specific time window. The specific time window can be set according to different risk scenarios and business needs. For example, for some short-term malicious behavior of obtaining activity benefits, the time window can be set to a few minutes or a few hours; for some long-term accumulated abnormal transaction behavior, the time window can be set to a day or a few days. Taking a set of combined risk rules targeting the malicious behavior of registering accounts in batches to obtain gift packs as an example, it may include dimensions such as a large number of newly registered accounts in a short period of time (e.g., within 1 hour) (user dimension), these accounts using the same or similar devices (device dimension), and these accounts participating in the same activity and obtaining a large number of gift packs (order and marketing activity dimension) and their sub-conditions. During the iterative matching process, if the relevant information in the judgment context satisfies the sub-conditions of all the dimensions in this set of combined risk rules, and these sub-conditions are simultaneously met within the set specific time window, then the combined risk rule is determined to be hit.
[0082] When a combined risk rule is triggered, a preliminary weighted risk score is calculated and aggregated based on the pre-configured weights of that rule. Different combined risk rules can be assigned different weights according to their impact on the game ecosystem and platform interests. For example, combined risk rules targeting serious disruptions to the game's economic order, such as recharge-for-cash-out activities, can be assigned higher weights, while those targeting less serious abnormal behaviors can be assigned lower weights. The preliminary weighted risk score is obtained by comprehensively calculating the weights of each triggered combined risk rule. This preliminary weighted risk score is then compared with a preset risk level threshold. The preset risk level threshold can be set according to the actual needs of game operation and risk tolerance. For example, it can be divided into three levels: low risk, medium risk, and high risk, with corresponding threshold ranges for each. Based on the comparison results, the corresponding primary risk marker and its level score are triggered. For example, if the preliminary weighted risk score falls within the high-risk level threshold range, a high-level primary risk marker is triggered, and a corresponding high-level score is assigned; if it falls within the medium-risk level threshold range, a medium-level primary risk marker and score are triggered; and if it falls within the low-risk level threshold range, a low-level primary risk marker and score are triggered. This approach enables timely and accurate identification of risky transactions, providing a basis for subsequent risk management.
[0083] In some embodiments, step S103 above, which involves inputting the primary risk marker into the trust scoring model and the anomaly detection model, and outputting a comprehensive risk score and handling recommendations by fusing the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the primary risk marker, specifically includes: Obtain the full-dimensional feature vector of the user corresponding to the primary risk label from the feature repository, and extract the user's long-term behavior feature subset, transaction real-time context and group behavior feature subset from the full-dimensional feature vector; A subset of users' long-term behavioral characteristics is input into a pre-trained trust scoring model. By analyzing the stability and value of users' historical behavior, a dynamic trust value is output. Input the real-time context of the transaction and a subset of group behavior features into a pre-trained anomaly detection model, and output anomaly scores by identifying the degree of deviation in the current transaction chain; The dynamic trust value, anomaly score, and the level score corresponding to the primary risk label are standardized and concatenated to form a fusion feature vector, which is then input into the integrated learning module for nonlinear relationship judgment, and outputs a comprehensive risk score and corresponding handling suggestions.
[0084] In this embodiment, when the system triggers a primary risk marker, it accurately locates and retrieves the full-dimensional feature vector of the user corresponding to the primary risk marker from the pre-built feature repository. This full-dimensional feature vector contains information about various aspects of the user's behavior and transactions within the game, such as the user's login time distribution, historical transaction frequency, and types of activities participated in. From the retrieved full-dimensional feature vector, three key subsets are further extracted. The first is a subset of long-term user behavior features, which focuses on the user's stable behavioral patterns over a long period, such as the user's average daily online time over the past few months or even years, and specific game play styles that the user has participated in for a long time. These features can reflect the stability and regularity of user behavior and help assess the user's credibility. The second is the real-time context of the transaction, which covers various real-time information at the time the current transaction occurs, such as the time point of the transaction, the ongoing marketing activities, and the status of the device used for the transaction. This information is crucial for judging the rationality of the current transaction. Third is the subset of group behavior characteristics. This subset considers the behavioral characteristics of other users in similar groups (such as the same game server, the same level range, etc.). By comparing group behavior, abnormal deviations in the current user's behavior can be discovered.
[0085] The extracted subset of long-term user behavioral features is input into a trust scoring model pre-trained on a large amount of data. During training, this model learns the long-term behavioral patterns of numerous normal and abnormal users, assessing a user's trustworthiness by analyzing the stability and value of their historical behavior. For example, if a user consistently logs into the game daily, participates in normal game activities, and their transaction behavior is consistent with established patterns, the model considers this user to have high trustworthiness and outputs a high dynamic trust value. Conversely, if a user's behavior frequently fluctuates abnormally, such as suddenly engaging in a large number of unusual transactions within a short period, the model outputs a low dynamic trust value. In this way, the trust scoring model can generate a dynamic trust assessment value for each user based on their long-term behavioral characteristics.
[0086] The real-time context of the transaction and a subset of group behavioral features are input into a pre-trained anomaly detection model. During training, the anomaly detection model learns various features and patterns of normal transaction chains, enabling it to identify the degree of deviation between the current transaction chain and the normal pattern. For example, if the current transaction occurs at an unusual time, uses a device different from the user's historically used device, and the transaction behavior differs significantly from the normal behavior of other users in the same group, then the anomaly detection model will consider this transaction to have a high anomaly risk and output a high anomaly score; conversely, if the transaction chain closely matches the normal pattern, it will output a lower anomaly score. In this way, the anomaly detection model can accurately determine the degree of anomaly of the current transaction.
[0087] The dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score corresponding to the primary risk label are standardized. The purpose of standardization is to unify data of different dimensions and ranges onto a common scale for subsequent fusion processing. The standardized dynamic trust value, anomaly score, and primary risk label level score are concatenated to form a fusion feature vector. This fusion feature vector integrates information from multiple aspects, including long-term user behavior, current transaction anomalies, and primary risk labels. The fusion feature vector is input into the ensemble learning module, which, by integrating the decision results of multiple base learners, can make non-linear relationship judgments, fully considering the complex relationships between various factors. After processing by the ensemble learning module, a comprehensive risk score is finally output, which comprehensively and accurately reflects the risk level of the current transaction. Simultaneously, based on the comprehensive risk score, the ensemble learning module also provides corresponding handling suggestions. For example, for high-risk transactions, it recommends immediate account freezing and transaction restrictions; for medium-risk transactions, it recommends further manual review or sending warning messages; and for low-risk transactions, normal processing but continuous monitoring is possible. In this way, targeted measures can be taken according to different risk situations, effectively ensuring the safe operation of game item gift packs and protecting the interests of the platform.
[0088] Furthermore, the step of inputting a subset of long-term user behavior features into a pre-trained trust scoring model, and outputting a dynamic trust value by analyzing the stability and value of the user's historical behavior, specifically includes: Construct a multi-task learning neural network model, which includes a shared low-level feature embedding layer and multiple parallel output heads. The shared low-level feature embedding layer is used to extract shared information from multiple features, and the parallel output heads are used to output the probability of a user engaging in fraudulent behavior in the future, the likelihood of long-term user retention, and the user's potential lifetime value level. A training sample set is constructed using historical user data, and a benchmark trust label is built based on users' continuous retention and paid performance within a preset time period. The training sample set and the baseline trust score label are used to train a multi-task learning neural network model. A subset of long-term user behavior features is input into the trained multi-task learning neural network model for prediction, and a preliminary static trust score is output. The system monitors the user's latest key behavioral events in real time, fine-tunes the initial static trust score based on a lightweight incremental calculation model, and introduces a time decay function to smoothly reduce the impact of long-standing historical behaviors, outputting the user's trust value.
[0089] In this embodiment, a multi-task learning neural network model is constructed, comprising a shared underlying feature embedding layer and multiple parallel output heads. The shared underlying feature embedding layer is designed to extract shared information from a subset of long-term user behavior features. This shared information can capture common patterns and regularities in user behavior, providing foundational support for subsequent analysis by the parallel output heads. For example, it can extract common feature information such as the frequency of user logins and the frequency of participation in different types of activities. The multiple parallel output heads each undertake different tasks. One output head outputs the probability of a user engaging in fraudulent behavior in the future, predicting the likelihood of future fraudulent operations by analyzing the user's historical behavior patterns. Another output head outputs the likelihood of long-term user retention, determining the probability of the user continuing to play the game in the future based on past behavior. A third output head outputs the user's potential lifetime value level, assessing the value the user may bring to the platform throughout the game's lifecycle. Through this multi-task learning architecture, the model can comprehensively utilize various aspects of user behavior information to more fully assess the user's trust status.
[0090] Historical user data is used to construct the training sample set. This data contains a wealth of user behavior information over a period of time, such as login time, transaction records, and activity participation. Simultaneously, baseline trust labels are constructed based on users' sustained retention and spending performance within a preset time period. For example, a three-month preset time period is set; users who remain engaged and make a certain amount of spending within these three months are assigned a higher baseline trust label, while users who churn or make very little spending within three months are assigned a lower baseline trust label. In this way, each training sample is labeled with a corresponding trust label, providing a clear reference standard for subsequent model training.
[0091] The constructed training sample set and corresponding baseline trust labels are input into a multi-task learning neural network model for training. During training, the model continuously adjusts its parameters to ensure that the outputs of each parallel output head are as close as possible to the true baseline trust labels. Through extensive data training, the model gradually learns the complex relationship between user behavior features and trust levels. After training, a subset of long-term user behavior features is input into the trained multi-task learning neural network model for prediction. Based on the input feature subset, the model calculates and outputs a preliminary static trust score through the calculations of each parallel output head. This preliminary static trust score is a relatively stable trust assessment value based on the user's past behavior, but it does not consider recent changes in the user's behavior.
[0092] To make trust assessment more real-time and accurate, it's necessary to monitor the user's latest key behavioral event stream in real time. These key behavioral event streams include the user's recent login behavior, transaction behavior, and activity participation. Based on a lightweight incremental calculation model, the initial static trust score is fine-tuned in real time according to these latest behavioral events. For example, if a user recently made a large, legitimate transaction, their trust score can be appropriately increased; conversely, if the user exhibits abnormal behavior, such as frequently changing login devices, their trust score can be appropriately decreased. Simultaneously, a time decay function is introduced, where the influence of long-term historical behavior on the current trust score gradually decreases over time. This is because user behavior patterns may change over time, and over-reliance on long-term historical behavior can lead to inaccurate trust assessments. The time decay function smoothly reduces the influence of long-term historical behavior, making the trust score more reflective of the user's current behavior. After real-time fine-tuning and time decay processing, the final output user trust value is the dynamic trust value, which more accurately reflects the user's trust status at the current moment.
[0093] Furthermore, the step of inputting a subset of the transaction's real-time context and group behavior features into a pre-trained anomaly detection model, and outputting anomaly scores by identifying the degree of deviation in the current transaction chain, specifically includes: A baseline model is trained using historical normal transaction data, and transaction context is introduced as a condition variable during the training process to enable the baseline model to learn the differences in normal behavior patterns under different scenarios, thus obtaining a dynamic baseline model. The feature combination of the real-time context of the transaction and a subset of group behavior features is input into the dynamic baseline model. The original anomaly index is generated by calculating the degree of deviation between the feature combination and the built-in normal pattern of the model. The original anomaly indicators are standardized to output a quantitative anomaly score that characterizes the probability of anomalies in the current trading behavior.
[0094] In this embodiment, a large amount of historical normal transaction data is selected as the training dataset. This historical normal transaction data covers rich information under various normal transaction scenarios during game operation, such as transaction records corresponding to different time periods, different user groups, and different types of item packs. During the training of the baseline model, transaction context is introduced as a condition variable. Transaction context includes various factors, such as the time of the transaction (e.g., during a game event or during regular hours); the type of game item pack involved in the transaction (different packs have different values and popularity, resulting in different normal transaction patterns); the account level of the user initiating the transaction (high-level and low-level accounts may have different transaction behavior patterns); and the device information used for the transaction (different devices may correspond to different user habits). By incorporating these transaction contexts as condition variables into the training process of the baseline model, the baseline model can learn the differences in normal behavior patterns under different scenarios. For example, during large-scale game events, the normal transaction frequency of certain popular packs will be significantly higher than during regular hours, and the model can learn this difference and form corresponding normal behavior patterns. After training and adjustment with a large amount of data, a dynamic baseline model is finally obtained. This model can dynamically determine whether transaction behavior conforms to a normal pattern based on different transaction contexts.
[0095] When a new transaction occurs, the system collects the immediate transaction context and a subset of group behavior features. The immediate transaction context includes the specific time of the transaction, the item bundle involved, and the current status of the account initiating the transaction. The group behavior feature subset reflects the recent behavioral characteristics of a group of users related to the transaction, such as the average purchase frequency and average spending amount of other users who purchased the item bundle within the same time period. These features are combined to form a complete feature vector, which is then input into a pre-trained dynamic baseline model. The dynamic baseline model calculates the deviation between the input feature combination and the model's built-in normal patterns based on the normal behavior patterns it has learned in different scenarios. For example, if the transaction occurs during an inactive period, but the transaction amount is significantly higher than the normal average transaction amount for the item bundle during inactive periods, and the account level initiating the transaction does not match the amount, the model will determine that the feature combination deviates significantly from the normal pattern and generate a raw anomaly indicator. This raw anomaly indicator is a preliminary, unprocessed value used to represent the degree of deviation between the current transaction behavior and the normal pattern.
[0096] To make anomaly indicators more comparable and interpretable, they need to be standardized. Standardization can employ common methods, such as mapping the original anomaly indicators to a specific numerical range, like 0 to 1. Through standardization, the original anomaly indicators are converted into a quantitative anomaly score that characterizes the probability of anomalies in current trading behavior. For example, after standardization, a score closer to 1 indicates a higher probability of anomalies in current trading behavior, while a score closer to 0 indicates that current trading behavior is closer to a normal pattern. This quantitative anomaly score can intuitively reflect whether anomalies exist in the current trading and the degree of anomaly, providing important information for subsequent risk assessment and handling.
[0097] This invention inputs a subset of real-time transaction context and group behavior features into a pre-trained anomaly detection model and outputs anomaly scores, which can effectively identify the degree of deviation in the current transaction chain and provide strong support for the safe operation of game item gift packs.
[0098] In some embodiments, step S104 above, which involves using the results of the comprehensive risk score, the implementation of disposal recommendations, and their feedback as state input to a deep Q-network, and automatically adjusting the risk score threshold, the intensity of disposal actions, and the weights of combined risk rules based on real-time risk control effectiveness indicators, specifically includes: The comprehensive risk score, the execution actions of the disposal recommendations, and the user feedback data are set as state vectors, and the adjustment instructions for the risk control score threshold, the intensity of disposal actions, and the weight of the combined rules are set as action space. Design a multi-objective weighted reward function based on the safety effect and user experience after action execution; Train a deep Q-network based on the state vector, action space, and multi-objective weighted reward function to learn a mapping strategy from state to optimal parameters to adjust actions; The real-time risk control performance data is input into the trained deep Q network. Based on the highest expected long-term reward output by the deep Q network, an action is selected to generate specific risk control parameter adjustment instructions.
[0099] In this embodiment, multiple data points need to be collected for the state vector. The comprehensive risk score is a key indicator reflecting the current level of risk in a transaction or behavior. It integrates information from trust scoring models, anomaly detection models, and primary risk markers, providing a clear picture of the risk level. Execution actions for proposed actions, such as restricting account login, freezing transactions, or sending warning messages, have varying impacts on risk control effectiveness and user experience. User feedback data includes user acceptance of the proposed actions and whether they raised objections, reflecting the rationality and effectiveness of the actions. Integrating these comprehensive risk scores, proposed actions, and user feedback data into a state vector comprehensively covers key information about the current risk control status. The action space is a set of adjustment instructions for risk control score thresholds, action intensity, and combined rule weights. For example, the risk control score threshold can be adjusted upwards or downwards by a certain margin; the action intensity can be increased or decreased, such as the severity of warning messages or the duration of account restrictions; and the combined rule weights can be redistributed for combined rules of different dimensions. These adjustment instructions constitute the action space, providing a range for subsequent deep Q-network action selection.
[0100] When designing the reward function, it is necessary to comprehensively consider two important objectives: the security effect after the action is executed and the user experience. Regarding security effect, if the action successfully intercepts malicious behavior by black market organizations, reducing economic losses and ecological damage within the game, a positive reward should be given. The reward value can be determined based on the severity and scope of the interception. For example, successfully intercepting a large-scale monetization attempt through third-party top-ups would grant a higher positive reward. Conversely, if the action fails to effectively intercept malicious behavior and even leads to further spread of black market activities, a negative reward should be given. Regarding user experience, if the action is reasonable and user feedback is positive, without causing user dissatisfaction or complaints, a positive reward should be given; conversely, if the action is too forceful or unreasonable, leading to numerous user complaints, a negative reward should be given. To balance these two objectives, a multi-objective weighted approach is adopted. Based on the actual needs and priorities of game operation, different weights are set for security effect and user experience, and these are combined to form a multi-objective weighted reward function. This function can provide corresponding reward values based on the actual situation after the action is executed, guiding the deep Q-network to learn in a direction that both ensures security and optimizes user experience.
[0101] Based on the pre-defined state vectors, action space, and multi-objective weighted reward function, a deep Q-network is trained. During training, different state vectors are input into the deep Q-network, and the network selects an action from the action space to execute based on the current state. After executing the action, the network obtains the corresponding reward value according to the multi-objective weighted reward function, and simultaneously observes the new state after the action. By continuously repeating this process, the deep Q-network gradually learns which actions to choose in different states to obtain the highest long-term reward; that is, it learns a mapping strategy from state to optimal parameter adjustment actions. For example, in a certain state, if adjusting the risk control score threshold upwards by a certain amount while reducing the intensity of the action yields a higher reward value, the network will gradually tend to choose this action combination in that state. After extensive training data and iterative training, the deep Q-network can form a relatively accurate and stable mapping strategy, providing effective decision support for subsequent real-time risk control.
[0102] During actual game operation, real-time status data of risk control effectiveness metrics are collected, such as the number of successfully intercepted black market activities, user complaint rate, and transaction success rate of normal users within a certain period. This real-time status data is input into a pre-trained deep Q-network. The deep Q-network evaluates different actions based on its learned mapping strategies and calculates the expected long-term reward for each action. Then, it selects the action with the highest expected long-term reward and generates specific risk control parameter adjustment instructions based on this action. For example, if the deep Q-network chooses to adjust the weight of the combined rule, increasing the weight of the account behavior dimension, then it will generate corresponding adjustment instructions to increase the weight of account behavior-related elements in the combined risk rules. In this way, risk control parameters can be automatically and dynamically adjusted based on real-time risk control effectiveness metrics, improving the accuracy and effectiveness of risk control and better protecting the safe operation of game item packs and the platform's interests.
[0103] In some embodiments, step S105 above, which involves using SHAP model interpretability technology to perform attribution analysis on the model decision-making process of high-risk cases and outputting the contribution of key risk features, specifically includes: Filter out transaction cases that are judged to be high-risk, and obtain the first fused feature vector and the corresponding high-risk prediction result obtained by inputting the transaction case into the trust scoring model and the anomaly detection model; Using the SHAP interpretability framework and based on a representative background dataset, the contribution of each feature in the first fused feature vector to the Shapley value of the high-risk prediction result is calculated. The features are sorted according to the absolute value of their contribution to the Shapley value, key risk-driving features are identified, and the direction and magnitude of each feature's contribution are analyzed to obtain the analysis results. Based on the analysis results, a multimodal interpretability report is generated, which includes text summaries, visualization charts, and structured data.
[0104] In this embodiment, the game operation risk control system continuously assesses the risks of various transactions. Based on the aforementioned comprehensive risk scoring and other judgment criteria, it accurately filters out high-risk transactions from a large number of cases. These high-risk cases may involve malicious top-up services by black market organizations, illegal acquisition of gift packs, and other similar activities. For each selected high-risk transaction, it is necessary to obtain the first fusion feature vector obtained by inputting it into the trust scoring model and the anomaly detection model. This first fusion feature vector is formed by integrating multi-dimensional information such as account information, device information, network behavior information, and gift pack-related data, comprehensively reflecting the characteristics of the transaction. Simultaneously, it is also necessary to obtain the high-risk prediction result corresponding to the transaction. This result is derived by the trust scoring model and the anomaly detection model based on the input feature vector, indicating the degree and probability of the transaction being judged as high-risk.
[0105] The SHAP interpretability framework is used for subsequent analysis. SHAP is an effective tool for interpreting the prediction results of machine learning models, based on the Shapley value concept in cooperative game theory. To calculate the Shapley value contribution of each feature in the first fused feature vector to the high-risk prediction result, a representative background dataset is needed. This background dataset should cover typical feature data of various normal and abnormal transactions in game operation, representing the overall distribution of transaction data. Based on this representative background dataset, the algorithmic logic of the SHAP interpretability framework is used to calculate the Shapley value contribution of each feature in the first fused feature vector to the high-risk prediction result. The Shapley value contribution reflects the magnitude and direction of the role each feature plays in the model's high-risk prediction decision; a positive value indicates that the feature tends to lead the model to predict high risk, while a negative value indicates that it tends to lead the model to predict low risk.
[0106] After obtaining the Shapley value contribution of each feature, all features are ranked according to the absolute value of their Shapley value contribution. A larger absolute value indicates a greater impact of the feature on the high-risk prediction result. This ranking clearly identifies the key risk-driving features that have the most significant impact on the high-risk prediction result. For example, certain specific account behavior features, device anomaly features, or network behavior features may rank highly and become key risk-driving features. After identifying the key risk-driving features, the direction and magnitude of each feature's contribution are further analyzed. For each key risk-driving feature, it is determined whether it contributes positively (causing the model to classify it as high-risk) or negatively (suppressing the model's classification as high-risk), and the specific numerical value of its contribution. This analysis provides a deeper understanding of the specific mechanism by which each key risk-driving feature plays a role in the model's decision-making process, offering targeted support for subsequent risk prevention and control.
[0107] Based on the above analysis results, a multimodal interpretability report is generated, including text summaries, visualizations, and structured data. The text summary concisely outlines the key risk-driving features of high-risk cases, along with the direction and magnitude of each feature's contribution, allowing readers to quickly grasp the core aspects of the case. The visualizations use intuitive graphical methods to display the distribution and contribution comparisons of key risk-driving features. For example, bar charts can be used to show the absolute value of each feature's Shapley value contribution, and line charts can show the directional changes in feature contribution, making complex data easier to understand and interpret. The structured data section presents all features in the first fusion feature vector and their corresponding Shapley value contributions and rankings in a standardized tabular format, facilitating data querying and analysis. This multimodal interpretability report comprehensively and in detail demonstrates the model decision-making process and key risk feature contributions of high-risk cases to game operators and risk management personnel, providing strong support for further optimizing risk control strategies and strengthening game security operations.
[0108] This embodiment uses SHAP model interpretability technology to perform attribution analysis on the model decision-making process of high-risk cases and outputs the contribution of key risk characteristics, providing an interpretable and traceable risk control decision-making basis for the safe operation of game item gift packs and the protection of platform interests.
[0109] Reference Figure 2 An embodiment of the present invention provides an AI-based gift pack top-up identification and risk control system 2, wherein the system 2 specifically includes: The data acquisition module 201 is used to build a feature warehouse based on the gift pack lifecycle data and marketing data from the game server, combined with account, device and network behavior data collected by the security component. Rule configuration module 202 is used to configure a set of combined risk rules containing multi-dimensional factors based on the feature warehouse. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, a primary risk label is triggered. The risk scoring module 203 is used to input the primary risk marker into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the primary risk marker, it outputs a comprehensive risk score and disposal recommendations. The risk adjustment module 204 is used to take the results of the comprehensive risk score, the implementation and disposal suggestions and their feedback as the status input to the deep Q network, and automatically adjust the risk score threshold, the intensity of disposal actions and the weight of combined risk rules according to the real-time risk control effect indicators. The attribution analysis module 205 is used to perform attribution analysis on the model decision-making process of high-risk cases using SHAP model interpretability technology, and output the contribution of key risk features.
[0110] It is understandable that, such as Figure 1 The content of the AI-based gift pack top-up identification and risk control method embodiments shown herein is applicable to the AI-based gift pack top-up identification and risk control system embodiments. The specific functions implemented by the AI-based gift pack top-up identification and risk control system embodiments are as follows: Figure 1 The AI-based gift pack top-up identification and risk control method shown is the same as the example, and the beneficial effects achieved are the same as those described above. Figure 1 The AI-based gift package top-up identification and risk control method shown in the embodiment achieves the same beneficial effects.
[0111] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0113] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the AI-based gift package top-up identification and risk control method as described in any of the above methods.
[0114] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0115] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0116] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0117] This invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is run by a processor, it implements the AI-based gift package top-up identification and risk control method as described in any of the above methods.
[0118] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0119] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0120] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0121] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0122] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. An AI-based method for identifying and controlling the risk of third-party top-up services for gift packages, characterized in that, The method specifically includes: A feature repository is constructed based on the lifecycle data of gift packs and marketing data from game servers, combined with account, device, and network behavior data collected by security components. Based on the feature warehouse, a set of combined risk rules containing multi-dimensional factors are configured. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, a primary risk label is triggered. The initial risk markers are input into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the initial risk markers, a comprehensive risk score and handling recommendations are output. The results of comprehensive risk scoring, implementation and disposal recommendations, and their feedback are used as state inputs to a deep Q-network, which automatically adjusts the risk scoring threshold, disposal action intensity, and weights of combined risk rules based on real-time risk control effectiveness indicators. The SHAP model interpretability technique is used to perform attribution analysis on the model decision-making process of high-risk cases and output the contribution of key risk features.
2. The method according to claim 1, characterized in that, The gift pack lifecycle data and marketing data based on the game server, combined with account, device, and network behavior data collected by the security component, are used to construct a feature warehouse, specifically including: Through the application programming interface and data bus, it accesses gift pack transaction logs from the game business server, activity data from the marketing platform, and account, device, and network behavior data collected from the client security component, forming multi-source heterogeneous data; Multi-source heterogeneous data is formatted, cleaned, and standardized to form standardized data. Based on standardized data, users, devices, orders, and marketing activities are extracted as core entities, and the binding, purchase, participation, and asset transfer relationships between core entities are established through order and behavior logs; For core entities, statistical calculations are performed within a preset time sliding window to generate temporal behavioral features and aggregation environment features; Standardized data, relationships, temporal behavioral features, and clustering environment features are concatenated and encoded according to core entities to form a feature warehouse.
3. The method according to claim 1, characterized in that, The feature warehouse-based configuration includes a set of combined risk rules containing multi-dimensional factors. When multiple factors simultaneously meet abnormal conditions within a specific time window, a primary risk marker is triggered, specifically including: Based on the quantitative feature factors defined in the feature warehouse, configure multiple sets of combined risk rules; Upon receiving a transaction risk assessment request, the system queries and assembles the feature vectors in the feature repository corresponding to the transaction risk assessment request in real time to form a assessment context. The judgment context is iterated and matched against all combined risk rules. When the sub-conditions of all dimensions of any combined risk rule are met simultaneously within a specific time window, the combined risk rule is judged to be hit. Based on the pre-configured weights of the hit combined risk rules, a preliminary weighted risk score is calculated and summarized. By comparing the preliminary weighted risk score with the preset risk level threshold, the corresponding primary risk label and its level score are triggered.
4. The method according to claim 1, characterized in that, The process involves inputting the initial risk marker into the trust scoring model and the anomaly detection model. By fusing the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the initial risk marker, a comprehensive risk score and handling recommendations are output, specifically including: Obtain the full-dimensional feature vector of the user corresponding to the primary risk label from the feature repository, and extract the user's long-term behavior feature subset, transaction real-time context and group behavior feature subset from the full-dimensional feature vector; A subset of users' long-term behavioral characteristics is input into a pre-trained trust scoring model. By analyzing the stability and value of users' historical behavior, a dynamic trust value is output. Input the real-time context of the transaction and a subset of group behavior features into a pre-trained anomaly detection model, and output anomaly scores by identifying the degree of deviation in the current transaction chain; The dynamic trust value, anomaly score, and the level score corresponding to the primary risk label are standardized and concatenated to form a fusion feature vector, which is then input into the integrated learning module for nonlinear relationship judgment, and outputs a comprehensive risk score and corresponding handling suggestions.
5. The method according to claim 4, characterized in that, The step of inputting a subset of long-term user behavior features into a pre-trained trust scoring model, and outputting a dynamic trust value by analyzing the stability and value of the user's historical behavior, specifically includes: Construct a multi-task learning neural network model, which includes a shared low-level feature embedding layer and multiple parallel output heads. The shared low-level feature embedding layer is used to extract shared information from multiple features, and the parallel output heads are used to output the probability of a user engaging in fraudulent behavior in the future, the likelihood of long-term user retention, and the user's potential lifetime value level. A training sample set is constructed using historical user data, and a benchmark trust label is built based on users' continuous retention and paid performance within a preset time period. The training sample set and the baseline trust score label are used to train a multi-task learning neural network model. A subset of long-term user behavior features is input into the trained multi-task learning neural network model for prediction, and a preliminary static trust score is output. The system monitors the user's latest key behavioral events in real time, fine-tunes the initial static trust score based on a lightweight incremental calculation model, and introduces a time decay function to smoothly reduce the impact of long-standing historical behaviors, outputting the user's trust value.
6. The method according to claim 4, characterized in that, The process of inputting a subset of real-time transaction context and group behavior features into a pre-trained anomaly detection model, and outputting anomaly scores by identifying the degree of deviation in the current transaction chain, specifically includes: A baseline model is trained using historical normal transaction data, and transaction context is introduced as a condition variable during the training process to enable the baseline model to learn the differences in normal behavior patterns under different scenarios, thus obtaining a dynamic baseline model. The feature combination of the real-time context of the transaction and a subset of group behavior features is input into the dynamic baseline model. The original anomaly index is generated by calculating the degree of deviation between the feature combination and the built-in normal pattern of the model. The original anomaly indicators are standardized to output a quantitative anomaly score that characterizes the probability of anomalies in the current trading behavior.
7. The method according to claim 1, characterized in that, The process of using the comprehensive risk score, the results of the proposed actions, and their feedback as state inputs to a deep Q-network, and automatically adjusting the risk score threshold, the intensity of actions, and the weights of combined risk rules based on real-time risk control effectiveness indicators, specifically includes: The comprehensive risk score, the execution actions of the disposal recommendations, and the user feedback data are set as state vectors, and the adjustment instructions for the risk control score threshold, the intensity of disposal actions, and the weight of the combined rules are set as action space. Design a multi-objective weighted reward function based on the safety effect and user experience after action execution; Train a deep Q-network based on the state vector, action space, and multi-objective weighted reward function to learn a mapping strategy from state to optimal parameters to adjust actions; The real-time risk control performance data is input into the trained deep Q network. Based on the highest expected long-term reward output by the deep Q network, an action is selected to generate specific risk control parameter adjustment instructions.
8. The method according to any one of claims 1 to 7, characterized in that, The attribution analysis of the model decision-making process for high-risk cases using the SHAP model interpretability technique outputs the contribution of key risk features, specifically including: Filter out transaction cases that are judged to be high-risk, and obtain the first fused feature vector and the corresponding high-risk prediction result obtained by inputting the transaction case into the trust scoring model and the anomaly detection model; Using the SHAP interpretability framework and based on a representative background dataset, the contribution of each feature in the first fused feature vector to the Shapley value of the high-risk prediction result is calculated. The features are sorted according to the absolute value of their contribution to the Shapley value, key risk-driving features are identified, and the direction and magnitude of each feature's contribution are analyzed to obtain the analysis results. Based on the analysis results, a multimodal interpretability report is generated, which includes text summaries, visualization charts, and structured data.
9. An AI-based system for identifying and controlling the risk of unauthorized top-up of gift packages, characterized in that, The system specifically includes: The data acquisition module is used to build a feature repository based on the lifecycle data of gift packs and marketing data from the game server, combined with account, device and network behavior data collected by the security component. The rule configuration module is used to configure a set of combined risk rules containing multi-dimensional factors based on the feature warehouse. When the multi-dimensional factors simultaneously meet the abnormal conditions within a specific time window, a primary risk label is triggered. The risk scoring module is used to input the primary risk markers into the trust scoring model and the anomaly detection model. By integrating the dynamic trust value output by the trust scoring model, the anomaly score output by the anomaly detection model, and the level score of the primary risk markers, it outputs a comprehensive risk score and handling recommendations. The risk adjustment module is used to input the results of the comprehensive risk score, the implementation and disposal suggestions and their feedback into the deep Q network, and automatically adjust the risk score threshold, the intensity of disposal actions and the weight of combined risk rules according to the real-time risk control effect indicators. The attribution analysis module is used to perform attribution analysis on the model decision-making process of high-risk cases using SHAP model interpretability technology, and output the contribution of key risk features.
10. A computer device, characterized in that, include: The memory and processor, and the computer program stored in the memory, when the computer program is executed on the processor, implement the AI-based gift package top-up identification and risk control method as described in any one of claims 1 to 8.