Account risk prediction method and system based on machine learning
By building a data set of dynamic behavior characteristics of accounts and introducing a time-sequence convolution network and attention mechanism, combining local anomaly factor evaluation and multi-stage decision-making network, the problem of insufficient account risk identification capabilities in the existing technology is solved, and risk prediction with higher accuracy and lower misjudgment rates is achieved.
Patent Information
- Application Number
- CN202510925956.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-07
AI Technical Summary
When identifying account risks, the existing technology is inadequate in identifying account risks, especially critical risk accounts, which leads to a high misjudgment rate and fails to effectively capture the time dependence and context evolution relationship of account transaction behavior, making it difficult to cope with the gradual evolution of risk behavior.
A high-dimensional feature coding and context sequence analysis mechanism is used to construct an account dynamic behavior feature data set, a risk behavior evolution model is constructed by combining a time-series convolution network and attention mechanism, a response offset of the perturbation parameter monitoring model is introduced, a local anomaly factor evaluation mechanism and a multi-stage decision network is used for misjudgment boundary correction, and finally a multi-dimensional evaluation is carried out in combination with a risk control knowledge base.
It significantly improves the accuracy of characterizing abnormal behavior patterns, reduces the rate of misjudgment, enhances the ability to identify behavior gradual trends and sudden abnormalities, improves the accuracy of prediction and generalization of the model, and provides a robust intelligent risk control solution.
Smart Images

Figure CN120450708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an account risk prediction method and system based on machine learning. Background Art
[0002] Today's financial environment is increasingly complex, with increasingly diverse manifestations of account risk. User risky behavior can be highly concealed and deceptive. Machine learning-based account risk prediction methods have been widely applied in financial anti-fraud, risk control, and auditing scenarios. A common approach involves training classification models based on historical account transaction data and user behavioral characteristics to determine whether an account is potentially risky. However, existing technologies still face significant technical bottlenecks, particularly in the ability to identify critically risky accounts. These accounts often exhibit behavioral characteristics that fluctuate between normal and abnormal, neither violating established rules nor exhibiting obvious fraudulent characteristics. Due to a lack of fine-grained dynamic behavior modeling, they are often mistakenly classified as low-risk, making them potential targets for exploitation. Existing methods fail to fully consider the time-dependence and contextual evolution of account transaction behavior when constructing features, making them difficult to respond to the gradual evolution of risky behavior. For example, an account may gradually increase its transaction limit, extend its operation time, or change its trading time period over a short period of time. If these slow-moving changes are not effectively captured, the model can easily overlook the gradual accumulation of risk. Therefore, it is necessary to design an account risk prediction method and system based on machine learning to reduce the misjudgment rate. Summary of the Invention
[0003] (1) Technical problems solved In response to the shortcomings of the existing technology, the present invention provides an account risk prediction method and system based on machine learning, which has the advantage of reducing the misjudgment rate and solves the problems in the above-mentioned background technology.
[0004] (2) Technical solution To achieve the above-mentioned purpose of reducing the misjudgment rate, the present invention provides the following technical solution: a method for predicting account risk based on machine learning, comprising the following steps: Obtain the target account's transaction details, login behavior, and terminal environment information in different time periods, and combine high-dimensional feature encoding strategies with contextual sequence analysis mechanisms to construct an account dynamic behavior feature dataset; We perform multi-layer nested feature extraction on the account dynamic behavior feature dataset and build an account risk behavior evolution model using a fusion algorithm of a temporal convolutional network and an attention mechanism. We also introduce perturbation parameters during the model construction process to monitor the model's response offset to abnormal behavior features in real time. Based on the trend change of the response offset, determine whether the risk sensitivity of the account risk behavior evolution model is stable in the current training cycle. If stable, record the time series evolution nodes of the risk behavior and dynamically correct the behavior sequence in combination with the local abnormal factor assessment mechanism; Based on the corrected behavior sequence, a multi-stage decision network and counterfactual reasoning mechanism are applied to generate a prediction score matrix for potential risk accounts and perform misjudgment boundary correction on the prediction results. Based on the results of the misjudgment boundary correction and combined with the preset risk control knowledge base, a multi-dimensional fusion assessment of the current status of the account is performed to output the final risk prediction results.
[0005] Preferably, the process of constructing the account dynamic behavior feature dataset is as follows: Collect the original behavioral data of the target account in different time periods, format the original data and remove outliers; Behavioral data is categorized and coded based on feature types. Z-score standardization is used for numerical features, and a combination of one-hot encoding and target encoding is used for categorical features. By utilizing the context sequence analysis mechanism, the dependency relationship between previous and subsequent behaviors is introduced into each behavioral time segment, context label features are constructed, and an account dynamic behavior feature dataset is generated.
[0006] Preferably, the multi-layer nested feature extraction process for the account dynamic behavior feature dataset is as follows: After inputting the feature dataset, a nested structure of short-term and long-term behaviors is constructed based on the behavior time slices; In the nested structure, local statistical feature extraction and trend change capture are performed separately to form nested feature representations at different time scales; A screening mechanism based on the rate of change of feature entropy is introduced to filter out invalid features and retain a subset of features with high information density; High information density feature subsets are hierarchically combined to generate derived features.
[0007] Preferably, the process of constructing the account risk behavior evolution model using the temporal convolutional network and attention mechanism fusion algorithm is as follows: The nested feature matrix is input into the temporal convolutional network, and the short-term change pattern is extracted using the local perception mechanism; After the temporal convolution layer, a multi-head self-attention mechanism is introduced to model the long-term dependencies and abnormal jump patterns between account behaviors. The introduction of residual connection structure enhances the nonlinear expression ability of the model; Combining behavior type labels and timestamp information, a sequence of intermediate states of risk evolution is generated, and the risk change prediction value is output synchronously during the training process; During the training process, perturbation parameters are introduced to simulate the changes in model output under input perturbations, and ultimately form an account risk behavior evolution model.
[0008] Preferably, the process of determining whether the risk sensitivity of the account risk behavior evolution model is stable in the current training cycle is: During each round of model training, the predicted output offset before and after the disturbance is calculated in real time, and the corresponding disturbance response vector is recorded; Fit the changing trend of the perturbation response vector in multiple consecutive training rounds and construct a response offset trend curve; Set the offset volatility threshold and the average disturbance response amplitude threshold to determine whether the model's risk perception ability is stable; If the response deviation trend curve is stable and the average disturbance response amplitude fluctuation is less than or equal to the deviation volatility threshold, the risk perception ability of the model is determined to be stable; If the response deviation trend curve is not stable and the average disturbance response amplitude fluctuation is greater than the deviation volatility threshold, the risk perception ability of the model is judged to be unstable.
[0009] Preferably, the process of dynamically correcting the behavior sequence in combination with the local abnormal factor evaluation mechanism is as follows: Extract risk evolution nodes and corresponding behavior fragments during the construction process and construct node local feature subsets; Apply the local outlier factor algorithm to each node's local feature subset to calculate the local anomaly score; Identify behavioral segments with high abnormality scores and determine that the behavioral segments are mutation behaviors; Combined with historical behavior templates of similar accounts, abnormal segments are replaced or expanded through matching analysis to complete abnormal behavior correction.
[0010] Preferably, the process of correcting the misjudgment boundary of the prediction result is: Generate an account risk score vector based on the corrected behavior sequence, and build a false positive boundary discrimination model based on the false positive samples in the training set. Introducing a counterfactual reasoning mechanism to construct contrasting behavior samples for boundary behaviors near the predicted results and simulate the possibility of boundary reversal; Adjust the discrimination threshold in the risk scoring model based on the deviation between the actual misclassified label and the simulated inversion output; Perform probability weighting calculation on the account scoring results in the boundary area to correct the deviation caused by fuzzy judgment.
[0011] Preferably, the process of outputting the final risk prediction result is: The risk score results after the false positive boundary correction are integrated with the account's historical behavior patterns, environmental labels, and device trustworthiness. Call the risk level mapping rules preset in the risk control knowledge base and make risk classification judgments based on the scoring range; For accounts with risk levels in the critical range, auxiliary judgment is made by combining the period risk weight and cross-platform behavior comparison score to output risk prediction results.
[0012] An account risk prediction system based on machine learning, comprising: Data collection module: This module obtains multi-source data such as the target account's transactions, logins, and terminal environment, and combines high-dimensional feature encoding with context sequence analysis to construct a dataset of account dynamic behavior characteristics. Risk Evolution Module: This module extracts multi-layer features from account behavior data, models the evolution of account risk behavior through the fusion of a temporal convolutional network and an attention mechanism, and introduces perturbation parameters to monitor model response deviations. Trend judgment module: determines whether the risk sensitivity is stable based on the model response deviation trend. If it is stable, it identifies the key nodes of behavior evolution and dynamically corrects the behavior sequence using the local abnormal factor mechanism; Risk scoring module: This module evaluates the potential risks of the corrected behavior sequence, generates a prediction matrix by combining a multi-stage decision network with counterfactual reasoning, and corrects the margin of misjudgment. Result output module: Integrates the risk control knowledge base and account risk scoring results, conducts multi-dimensional fusion analysis, and outputs risk prediction results.
[0013] (3) Beneficial effects Compared with the existing technology, the present invention provides an account risk prediction method and system based on machine learning, which has the following beneficial effects: By introducing high-dimensional feature encoding and context sequence analysis mechanisms, the present invention can fully explore the behavioral evolution information of accounts in different time periods and significantly improve the accuracy of characterizing abnormal behavior patterns; it adopts a fusion of temporal convolutional networks and attention mechanisms to construct a risk behavior evolution model, taking into account global dependencies while ensuring local behavior perception, and enhancing the ability to identify gradual behavioral trends and sudden anomalies; by introducing perturbation parameters and real-time monitoring of model output offsets, it effectively judges the stability of risk sensitivity and avoids misjudgments caused by using unconverged models; it combines the local anomaly factor evaluation mechanism to correct the behavior sequence, further improving the credibility and consistency of the behavior data; the multi-stage decision network and counterfactual reasoning mechanism can dynamically correct the boundary prediction results, significantly reducing the risk of misjudgment caused by sample distribution offset; and finally, it integrates the risk control knowledge base for multi-dimensional evaluation, effectively improving the predictive interpretability and business adaptability of the prediction. The overall solution has higher prediction accuracy, lower error rate and stronger model generalization ability, providing the financial industry with a robust, efficient and real-time adaptable intelligent risk control solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the method of the present invention; Figure 2 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0016] Example 1: Please refer to Figure 1 As shown, an account risk prediction method based on machine learning according to an embodiment of the present invention includes the following steps: S1: Obtain the target account's transaction details, login behavior, and terminal environment information in different time periods, and combine high-dimensional feature encoding strategies with context sequence analysis mechanisms to construct an account dynamic behavior feature dataset.
[0017] The process of constructing the account dynamic behavior feature dataset in S1 is as follows: Collect the target account's original behavioral data, including transaction details, login time, IP address changes, device information, geographic location, transaction amount, transaction frequency, and operation path, over different time periods. Format the original data and remove outliers. Behavioral data is categorized and coded based on feature type. Numerical features are standardized using Z-score, while categorical features are standardized using a combination of one-hot encoding and target encoding. Numerical features are standardized to a mean of 0 and a standard deviation of 1, improving scale consistency across features. For categorical features, high-frequency categories are one-hot encoded, converting each category into a separate binary column. Target encoding is used for low-frequency categories, and the average risk probability corresponding to each category is calculated based on historical risk labels and mapped to a continuous value. Using the context sequence analysis mechanism, we introduce the dependency relationship between previous and subsequent behaviors in each behavioral time segment, construct contextual label features including behavior transfer paths, behavior jump probabilities, and operation time preferences, and generate a dataset of dynamic account behavior features. We divide the behavior sequence into time segments according to fixed time windows or behavior continuity rules. Behavioral events within each time segment are organized sequentially to form a behavior chain. We model the behavior type sequence of each account within the segment, such as login → transfer → balance inquiry → exit, and record the behavior transition matrix. We introduce a state transition probability matrix to represent the transition tendency between different types of behaviors. We analyze whether there are jumps between consecutive behaviors that do not conform to normal operating habits and calculate the probability of jumps occurring. We map the time of behavioral events to the time period of the day and calculate the time preference distribution of account operations. We integrate the behavior chain sequence, transition probability, jump indicator, and time period preference with the original encoding features through feature splicing to generate a complete time series context feature matrix. We then unify the high-dimensional encoding features, time segment labels, and contextual behavior dependency structure after the above processing. We construct structured behavior sequence samples with account ID + time segment as the primary key, and each sample records the complete behavioral characteristics of an account within a specific time segment.
[0018] S2: Perform multi-layer nested feature extraction on the account dynamic behavior feature dataset, and use the temporal convolutional network and attention mechanism fusion algorithm to build an account risk behavior evolution model. In addition, perturbation parameters are introduced during the model construction process to monitor the model's response offset to abnormal behavior features in real time.
[0019] The multi-layer nested feature extraction process for the account dynamic behavior feature dataset in S2 is as follows: After inputting the feature dataset, a nested structure of short-term and long-term behaviors is constructed based on the behavior time slices; Within the nested structure, local statistical feature extraction and trend change capture are performed separately to form nested feature representations at different time scales. The account dynamic behavior feature dataset is sorted in chronological order, forming a complete behavior sequence for each account. A short-term window is set to capture instantaneous behavioral feature fluctuations. A long-term window is set to model the account's behavior patterns on a daily or weekly cycle. The behavior sequence in the long-term window is used as the outer feature structure, and the short-term window sequence as the inner behavior stream, forming a sliding window + time pyramid nested expression in the model structure. Short-term and long-term windows are paired for each time node to ensure the synchronization of nested feature time series. Introduce a screening mechanism based on the rate of change of feature entropy to filter out invalid features and retain a subset of features with high information density; for each behavioral feature , calculate the information entropy under the full sample, the formula is: ; Where, is the probability of eigenvalue distribution; The time rate of change of information entropy is calculated for the characteristic value sequence of the same account in a continuous time period. The formula is: ; Where, Features The rate of change of information entropy between adjacent time windows t-1 to t, Features The information entropy within the current time window t, Features Information entropy within the previous time window t-1; Hierarchically combine high-information-density feature subsets to generate derived features such as behavioral change speed, amplitude, and frequency; among the retained feature subsets, combine them according to the three-dimensional structure of feature-time-account; construct interactive features, such as: transaction amount × login time period, terminal type × operation frequency, IP change frequency × geographical span, etc.; for a certain numerical feature within the sliding time window, calculate the change rate, record the maximum jump amplitude and direction of the behavioral features in each window, count the frequency of occurrence of key behaviors per unit time, and construct a frequency vector.
[0020] The process of constructing the account risk behavior evolution model using the temporal convolutional network and attention mechanism fusion algorithm in S2 is as follows: The nested feature matrix is input into the temporal convolutional network, and the short-term change pattern is extracted using the local perception mechanism; the multi-layer nested feature representation is constructed as a three-dimensional tensor input based on the account. , where B is the number of accounts, T is the time step length, and F is the feature dimension of each time step; a temporal convolutional network is used to construct a stacked structure consisting of multiple one-dimensional convolutional layers. Dilated convolution is used to expand the receptive field. Each convolutional layer sliding window covers multiple time steps and only uses past information for feature extraction to avoid future data leakage. The convolution operation outputs a feature tensor containing the local behavior change pattern. After the temporal convolutional layer, a multi-head self-attention mechanism is introduced to model the long-term dependencies and abnormal jump patterns between account behaviors. The core idea of the multi-head self-attention mechanism is to adaptively learn the weights between behavioral events in the time dimension. The convolutional network output is converted into three sets of representations: Q, K, and V, which represent the query vector at the current time step, the attention key vector, and the behavior value, respectively. Multiple independent attention heads are used to capture the behavioral dependency characteristics at different levels or dimensions. Finally, the outputs of each head are spliced and projected. The introduction of residual connection structure enhances the nonlinear expression ability of the model and improves training stability. Adding residual connection after each convolutional layer and attention module allows the original features to be directly propagated in subsequent layers to avoid gradient vanishing. It improves the network's sensitivity to feature changes and improves learning stability. It supports deep stacked network structures to better fit the nonlinear risk evolution process. Combining behavior type labels with timestamp information, a sequence of intermediate states of risk evolution is generated, and the training process simultaneously outputs risk change predictions. The type labels of behavioral events are encoded into vectors through the embedding layer and concatenated to the output of the attention module. A standardized timestamp vector is added to enhance the model's ability to model temporal dependencies. The concatenated result is input into a bidirectional GRU layer, which outputs a representation of the risk state at each time step. A fully connected layer with a sigmoid activation function is applied to each time step to output a risk score, representing the potential risk level of the account at that point in time. The training objective is to minimize the cross-entropy loss between the predicted score and the true label. During the training process, perturbation parameters are introduced to simulate the changes in model output under input perturbations, and ultimately a stable and reusable account risk behavior evolution model is formed.
[0021] S3: Based on the trend change of the response offset, determine whether the risk sensitivity of the account risk behavior evolution model in the current training cycle is stable. If stable, record the time series evolution nodes of the risk behavior and dynamically correct the behavior sequence in combination with the local abnormal factor evaluation mechanism.
[0022] The process of judging whether the risk sensitivity of the account risk behavior evolution model in the current training cycle is stable in S3 is: During each round of model training, the predicted output offset before and after the disturbance is calculated in real time, and the corresponding disturbance response vector is recorded. During each round of model training, a perturbation is introduced to the input account dynamic nested feature data to form a disturbance sample. The original sample and the perturbation sample are respectively input into the account risk behavior evolution model to obtain two sets of prediction results. The original prediction output is compared with the prediction output of the perturbation sample, and the prediction difference between the two is calculated. This difference is recorded as the disturbance response vector for the current round. After each round of training is completed, the disturbance response vector is appended to the disturbance response sequence. The changing trend of the disturbance response vector in multiple consecutive training rounds is fitted, and a response offset trend curve is constructed. Based on the disturbance response vectors recorded during multiple consecutive training rounds, a certain number of continuous vectors are selected to form a sliding window. The trend analysis algorithm is applied to fit the changing trend of the disturbance response within the training cycle to obtain a disturbance offset trend curve. By analyzing the slope and fluctuation trend of the trend curve, the model's ability to respond stably to disturbance input during training is judged. Set the offset volatility threshold and the average disturbance response amplitude threshold to determine whether the model's risk perception ability is stable; calculate the standard deviation and mean of the disturbance response vector within the sliding window and compare them with the two set thresholds respectively; If the response deviation trend curve is stable and the average disturbance response amplitude fluctuation is less than or equal to the deviation volatility threshold, the risk perception ability of the model is determined to be stable; If the response deviation trend curve is not stable and the average disturbance response amplitude fluctuation is greater than the deviation volatility threshold, the risk perception ability of the model is judged to be unstable.
[0023] The process of dynamically correcting the behavior sequence in S3 by combining the local anomaly factor evaluation mechanism is as follows: Extract risk evolution nodes and corresponding behavior segments during the construction process to construct node local feature subsets. During the account risk behavior evolution model construction process, record the risk evolution nodes identified as key changes at each time step, and extract several behavior records before and after them to form local behavior segments. For each node behavior segment, extract the corresponding feature vector set, including transaction amount change rate, device switching frequency, login location jump amplitude, etc., to construct a node local feature subset for anomaly detection. Apply the local outlier factor algorithm to each node's local feature subset to calculate the local anomaly score. For each local feature subset, the local outlier factor algorithm is used to process it. The difference in K-nearest neighbor density between the segment and its neighboring segments is calculated to assess the degree of isolation of the behavior segment in the local feature space. An outlier factor score is output, with a higher score indicating a higher likelihood of local abnormal behavior. Identify behavioral segments with high anomaly scores and determine them as mutation behaviors. Filter all local outlier factor scores based on a preset anomaly score threshold. If the LOF value of a behavior segment exceeds the threshold and there is a clear temporal / spatial / pattern mutation in the original behavior sequence, the behavior segment is marked as mutation behavior. Mutation behaviors include abnormal transaction times, sudden transaction peaks, and rare device logins. Combined with historical behavior templates of similar accounts, abnormal segments are replaced or expanded through matching analysis to complete abnormal behavior correction.
[0024] It is understandable that the role of judging whether the risk sensitivity of the account risk behavior evolution model is stable within the current training cycle is: Function 1: By analyzing the trend changes of the perturbation response vector in consecutive training rounds, it can be used to identify whether the model's prediction results fluctuate significantly under different input perturbations. If the fluctuation is too large, it indicates that the model is overly sensitive to local abnormal behavior or unstable, and there is a risk of overfitting or structural underfitting. Function 2: When the risk sensitivity of the model tends to be stable during the training cycle, the identified risk evolution nodes and prediction results have a high degree of credibility, which can effectively support subsequent steps such as behavioral sequence anomaly correction and risk scoring matrix construction. If the stability is not judged and the output of the non-converged or unstable model is directly used, it is easy to cause the subsequent misjudgment boundary to expand, affecting the accuracy of the overall risk prediction chain.
[0025] The technical solution of this embodiment is: by introducing perturbation parameters during the model training process, monitoring the offset response of the model output before and after the perturbation in real time, and constructing the response offset trend curve under continuous rounds; if the fluctuation degree of the curve is lower than the preset threshold, it is determined that the model's perception of risky behaviors in the current training cycle is stable. Subsequently, the time-series evolution nodes of the risky behaviors identified by the model in the cycle are extracted, and the abnormal behavior fragments are identified and corrected based on the local anomaly factor evaluation mechanism, the account behavior sequence is dynamically corrected, and the subsequent risk scoring results are optimized. Effective perception of the risk sensitivity state of the model is achieved, avoiding premature reliance on its output results when the model is not yet stable, and improving the reliability of the risk identification results; at the same time, the behavior sequence is dynamically corrected through the local anomaly factor evaluation mechanism, significantly improving the model's fault tolerance and robustness to behavioral mutations, thereby reducing the misjudgment rate and enhancing prediction accuracy.
[0026] Example 2: Figure 1 As shown, a method for predicting account risk based on machine learning further includes the following steps: S4: Based on the corrected behavior sequence, a multi-stage decision network and counterfactual reasoning mechanism are applied to generate a prediction score matrix for potential risk accounts and perform misjudgment boundary correction on the prediction results.
[0027] The process of correcting the misjudgment boundary of the prediction result in S4 is as follows: Generate an account risk score vector based on the corrected behavior sequence, and build a misjudgment boundary discrimination model based on misjudgment samples from the training set. Use the anomaly-corrected account behavior sequence as model input and, through the trained risk behavior evolution model, output the risk score vector for the current account. Extract historical misjudgment samples from the training set, including instances that were incorrectly identified as risky accounts or normal accounts, and analyze their score distribution and behavioral characteristics. Input the normal, risk, and misjudgment samples into the boundary discrimination sub-model to train a boundary recognition module with misjudgment sensitivity to identify accounts prone to misjudgment within the scoring boundary area. A counterfactual reasoning mechanism is introduced to construct comparative behavioral samples for boundary behaviors near the predicted results to simulate the possibility of boundary reversal. A behavioral sequence whose risk score of the current account falls near the model's score boundary is selected, and a set of counterfactual behavioral samples is generated based on this sequence. This involves artificially controlling a small number of feature changes to keep the majority of the behavioral background unchanged. The original sequence and the counterfactual samples are input into the model together to observe whether the predicted labels are reversed. The probability of label change for each set of boundary samples under different perturbation conditions is calculated to measure the stability and boundary sensitivity of the judgment results. Adjust the discrimination threshold in the risk scoring model based on the deviation between the actual misjudgment label and the simulated inversion output; compare the true label of the historical misjudgment sample with the predicted result of its counterfactual simulated sample, and calculate the label deviation between the two in the model scoring space; if the misjudgment sample's label is easily flipped under the counterfactual perturbation, it indicates that the current model boundary setting has a sensitive interval; based on the error distribution and boundary behavior characteristics, dynamically adjust the risk discrimination score threshold in the model to have a higher error tolerance rate in the high misjudgment area, thereby improving the recognition accuracy of boundary samples; The account scoring results in the boundary area are probability-weighted to correct the offset caused by fuzzy judgment; for all accounts whose scores fall near the risk judgment boundary, their label uncertainty coefficients are calculated based on the reversal probability under the counterfactual simulation; the original risk score is probability-weighted using the uncertainty coefficient to improve the prediction value's tolerance for fuzzy samples; the weighted scoring results can dynamically reflect the risk fluctuation degree of account behavior under boundary disturbances, thereby forming a more robust final judgment result.
[0028] S5: Based on the error boundary correction results and combined with the preset risk control knowledge base, a multi-dimensional fusion assessment of the current status of the account is performed to output the final risk prediction results.
[0029] The process of outputting the final risk prediction result in S5 is as follows: The risk score results after the error boundary correction are integrated with the account's historical behavior patterns, environmental tags, and device trust. The account risk score results after the error boundary correction are obtained, and the historical behavior pattern characteristics of the target account over a long period of time are retrieved, including behavior rhythm, capital flow path, operating habits, etc. The environmental tags of the time period to which the current behavior belongs, such as holidays, nighttime visits, sensitive area logins, and other contextual factors, are extracted. The historical behavior records and associated risk distribution of the terminal device used by the account are obtained, and the trust score of the device is calculated. The above three types of information are used as weighted features and integrated with the risk score results. The integrated risk probability value is obtained through feature normalization and Bayesian inference. The system calls the risk level mapping rules preset in the risk control knowledge base and combines them with the scoring intervals to make a risk classification judgment. The integrated risk probability value is input into the risk level mapping function predefined in the risk control knowledge base. The function divides the probability value into several scoring intervals based on the standards of financial institutions. The system then assigns a preliminary risk level mark to the account based on the range of the integrated score. For accounts with risk levels in the critical interval, auxiliary judgment is made by combining the time period risk weight and cross-platform behavior comparison score to output the risk prediction result; for accounts on the critical edge of multiple risk levels, the risk weight parameter of the time period to which their behavior belongs is extracted, which represents the fraud probability of a specific time period based on historical statistics; the behavior trajectory of the account on other platforms or systems is analyzed for consistency with the behavior on the current platform, and the cross-platform behavior comparison score is calculated; the above two indicators are used to adjust the risk level label of the current account. If both indicators tend to be high risk, the critical account will be marked as high risk, otherwise the original level will be maintained; and finally, the risk prediction result after fusion judgment is output.
[0030] Example 3: Please refer to Figure 2 As shown, an account risk prediction system based on machine learning includes: Data collection module: This module obtains multi-source data such as the target account's transactions, logins, and terminal environment, and combines high-dimensional feature encoding with context sequence analysis to construct a dataset of account dynamic behavior characteristics. Risk Evolution Module: This module extracts multi-layer features from account behavior data, models the evolution of account risk behavior through the fusion of a temporal convolutional network and an attention mechanism, and introduces perturbation parameters to monitor model response deviations. Trend judgment module: determines whether the risk sensitivity is stable based on the model response deviation trend. If it is stable, it identifies the key nodes of behavior evolution and dynamically corrects the behavior sequence using the local abnormal factor mechanism; Risk scoring module: This module evaluates the potential risks of the corrected behavior sequence, generates a prediction matrix by combining a multi-stage decision network with counterfactual reasoning, and corrects the margin of misjudgment. Result output module: Integrates the risk control knowledge base and account risk scoring results, conducts multi-dimensional fusion analysis, and outputs risk prediction results.
[0031] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0032] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting account risk based on machine learning, characterized in that: The following steps are involved: Obtain the target account's transaction details, login behavior, and terminal environment information in different time periods, and combine high-dimensional feature encoding strategies with contextual sequence analysis mechanisms to construct an account dynamic behavior feature dataset; We perform multi-layer nested feature extraction on the account dynamic behavior feature dataset and build an account risk behavior evolution model using a fusion algorithm of a temporal convolutional network and an attention mechanism. We also introduce perturbation parameters during the model construction process to monitor the model's response offset to abnormal behavior features in real time. Based on the trend change of the response offset, determine whether the risk sensitivity of the account risk behavior evolution model is stable in the current training cycle. If stable, record the time series evolution nodes of the risk behavior and dynamically correct the behavior sequence in combination with the local abnormal factor assessment mechanism; Based on the corrected behavior sequence, a multi-stage decision network and counterfactual reasoning mechanism are applied to generate a prediction score matrix for potential risk accounts and perform misjudgment boundary correction on the prediction results. Based on the results of the misjudgment boundary correction and combined with the preset risk control knowledge base, a multi-dimensional fusion assessment of the current status of the account is performed to output the final risk prediction results.
2. The account risk prediction method based on machine learning according to claim 1, characterized in that: The process of constructing the account dynamic behavior feature dataset is as follows: Collect the original behavioral data of the target account in different time periods, format the original data and remove outliers; Behavioral data is categorized and coded based on feature types. Z-score standardization is used for numerical features, and a combination of one-hot encoding and target encoding is used for categorical features. By utilizing the context sequence analysis mechanism, the dependency relationship between previous and subsequent behaviors is introduced into each behavioral time segment, context label features are constructed, and an account dynamic behavior feature dataset is generated.
3. The account risk prediction method based on machine learning according to claim 2, characterized in that: The multi-layer nested feature extraction process for the account dynamic behavior feature dataset is as follows: After inputting the feature dataset, a nested structure of short-term and long-term behaviors is constructed based on the behavior time slices; In the nested structure, local statistical feature extraction and trend change capture are performed separately to form nested feature representations at different time scales; A screening mechanism based on the rate of change of feature entropy is introduced to filter out invalid features and retain a subset of features with high information density; High information density feature subsets are hierarchically combined to generate derived features.
4. The account risk prediction method based on machine learning according to claim 3, characterized in that: The process of constructing the account risk behavior evolution model using the fusion algorithm of temporal convolutional network and attention mechanism is as follows: The nested feature matrix is input into the temporal convolutional network, and the short-term change pattern is extracted using the local perception mechanism; After the temporal convolution layer, a multi-head self-attention mechanism is introduced to model the long-term dependencies and abnormal jump patterns between account behaviors. The introduction of residual connection structure enhances the nonlinear expression ability of the model; Combining behavior type labels and timestamp information, a sequence of intermediate states of risk evolution is generated, and the risk change prediction value is output synchronously during the training process; During the training process, perturbation parameters are introduced to simulate the changes in model output under input perturbations, and ultimately form an account risk behavior evolution model.
5. The account risk prediction method based on machine learning according to claim 4, characterized in that: The process of judging whether the risk sensitivity of the account risk behavior evolution model is stable in the current training cycle is as follows: During each round of model training, the predicted output offset before and after the disturbance is calculated in real time, and the corresponding disturbance response vector is recorded; Fit the changing trend of the perturbation response vector in multiple consecutive training rounds and construct a response offset trend curve; Set the offset volatility threshold and the average disturbance response amplitude threshold to determine whether the model's risk perception ability is stable; If the response deviation trend curve is stable and the average disturbance response amplitude fluctuation is less than or equal to the deviation volatility threshold, the risk perception ability of the model is determined to be stable; If the response deviation trend curve is not stable and the average disturbance response amplitude fluctuation is greater than the deviation volatility threshold, the risk perception ability of the model is judged to be unstable.
6. The account risk prediction method based on machine learning according to claim 5, characterized in that: The process of dynamically correcting the behavior sequence by combining the local anomaly factor evaluation mechanism is as follows: Extract risk evolution nodes and corresponding behavior fragments during the construction process and construct node local feature subsets; Apply the local outlier factor algorithm to each node's local feature subset to calculate the local anomaly score; Identify behavioral segments with high abnormality scores and determine that the behavioral segments are mutation behaviors; Combined with historical behavior templates of similar accounts, abnormal segments are replaced or expanded through matching analysis to complete abnormal behavior correction.
7. The account risk prediction method based on machine learning according to claim 6, characterized in that: The process of correcting the misjudgment boundary of the prediction results is as follows: Generate an account risk score vector based on the corrected behavior sequence, and build a false positive boundary discrimination model based on the false positive samples in the training set. Introducing a counterfactual reasoning mechanism to construct contrasting behavior samples for boundary behaviors near the predicted results and simulate the possibility of boundary reversal; Adjust the discrimination threshold in the risk scoring model based on the deviation between the actual misclassified label and the simulated inversion output; Perform probability weighting calculation on the account scoring results in the boundary area to correct the deviation caused by fuzzy judgment.
8. The account risk prediction method based on machine learning according to claim 7, characterized in that: The process of outputting the final risk prediction results is as follows: The risk score results after the false positive boundary correction are integrated with the account's historical behavior patterns, environmental labels, and device trustworthiness. Call the risk level mapping rules preset in the risk control knowledge base and make risk classification judgments based on the scoring range; For accounts with risk levels in the critical range, auxiliary judgment is made by combining the period risk weight and cross-platform behavior comparison score to output risk prediction results.
9. An account risk prediction system based on machine learning, applied to the method according to any one of claims 1 to 8, characterized in that: include: Data collection module: This module obtains multi-source data such as the target account's transactions, logins, and terminal environment, and combines high-dimensional feature encoding with context sequence analysis to construct a dataset of account dynamic behavior characteristics. Risk Evolution Module: This module extracts multi-layer features from account behavior data, models the evolution of account risk behavior through the fusion of a temporal convolutional network and an attention mechanism, and introduces perturbation parameters to monitor model response deviations. Trend judgment module: determines whether the risk sensitivity is stable based on the model response deviation trend. If it is stable, it identifies the key nodes of behavior evolution and dynamically corrects the behavior sequence using the local abnormal factor mechanism; Risk scoring module: This module evaluates the potential risks of the corrected behavior sequence, generates a prediction matrix by combining a multi-stage decision network with counterfactual reasoning, and corrects the margin of misjudgment. Result output module: Integrates the risk control knowledge base and account risk scoring results, conducts multi-dimensional fusion analysis, and outputs risk prediction results.
Citation Information
Patent Citations
User risk determination method and device and server
CN117593107A
Bank fraud detection method based on time domain convolutional network and attention mechanism
CN118194113A
Financial transaction risk control method based on financial sequence generation technology
CN118333763A
Business fund embezzlement risk identification method and device, electronic equipment and storage medium
CN118822707A
Defense system security assessment method and system based on time sequence attention mechanism
CN119449466A
Cited By
Risk prediction method based on mobile terminal equipment
CN121052829A
Short message fraud early warning method and system based on abnormal behavior detection
CN121262579A
User behavior intelligent monitoring method and system based on time axis linkage
CN121723468A