An intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction
Through multi-dimensional correlation analysis and deep feature learning, combined with multi-layer Transformer attention network and TR-A2C deep reinforcement learning algorithm, unified modeling of multi-channel transaction data and dynamic optimization of risk control strategies are achieved, and inconsistent data standards and risk control strategies in the bank payment and settlement system are solved, and the accuracy and real-time nature of risk control decisions are improved.
Patent Information
- Application Number
- CN202510371481.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Bank payment and settlement systems are facing challenges such as a surge in multi-channel transaction data, improved real-time risk control requirements, and bottlenecks in system processing performance. Traditional systems have problems such as inconsistent channel data standards, solidified risk control rules, and delayed system response.
Through multi-dimensional correlation analysis and deep feature learning, multi-channel transaction behavior data are uniformly modeled and feature extraction, and multi-layer Transformer attention network and TR-A2C deep reinforcement learning algorithm are used to realize dynamic optimization and adaptive adjustment of risk control strategies. Through a consistent hashing algorithm and load balancing strategy, distributed storage and dynamic scheduling of multi-channel data are realized.
It improves the accuracy and real-time nature of risk control decisions, solves the problems of system scalability and data access performance, ensures the consistency and reliability of multi-channel data, and effectively prevents the risk of transaction fraud.
Smart Images

Figure CN119887204B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of payment and settlement, and particularly relates to an intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction. Background Art
[0002] The bank payment and settlement system faces challenges such as a sharp increase in multi-channel transaction data, higher requirements for real-time risk control, and bottlenecks in system processing performance. Traditional payment and settlement systems generally have problems such as inconsistent channel data standards, fixed risk control rules, and system response delays, making it difficult to meet business requirements.
[0003] In the multi-channel payment scenario, there are significant differences in the transaction data formats and communication protocols of different channels, making data integration difficult and risk control rules difficult to uniformly manage and dynamically adjust. At the same time, with the growth of the transaction scale, the system needs to complete risk control decisions at the millisecond level, posing higher requirements for computing performance and processing delays. Summary of the Invention
[0004] The main object of the present invention is to provide an intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction. The present invention realizes the distributed storage and dynamic scheduling of multi-channel data, solves the problems of system scalability and data access performance, and improves the accuracy and real-time performance of risk control decisions.
[0005] To achieve the above object, the present invention provides an intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction, including the following steps:
[0006] Perform multi-dimensional correlation analysis on the multi-channel transaction behavior flow data to obtain transaction channel correlation data, and perform deep feature learning on the transaction channel correlation data to obtain transaction behavior feature vectors;
[0007] Perform protocol conversion and field mapping processing on the multi-channel original message data to obtain standard payment transaction data;
[0008] Perform stochastic network modeling on the multi-channel transaction processing flow according to the standard payment transaction data to obtain processing delay threshold parameters;
[0009] Construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain the multi-channel real-time risk control decision strategy through parallel policy optimization.
[0010] The present invention also provides an intelligent bank payment and settlement device supporting real-time risk control and multi-channel interaction, including:
[0011] An association analysis unit is used to perform multi-dimensional association analysis on multi-channel transaction behavior flow data to obtain transaction channel association data, and perform in-depth feature learning on the transaction channel association data to obtain transaction behavior feature vectors;
[0012] A conversion unit is used to perform protocol conversion and field mapping processing on multi-channel original message data to obtain standard payment transaction data;
[0013] A modeling and analysis unit is used to perform stochastic network modeling on the multi-channel transaction processing flow according to the standard payment transaction data to obtain a processing delay threshold parameter;
[0014] A risk control decision-making unit is used to construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain a multi-channel real-time risk control decision-making strategy through parallel policy optimization.
[0015] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0016] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0017] In summary, the technical solution provided by the present invention performs unified modeling and feature extraction on multi-channel transaction behavior data through multi-dimensional association analysis and in-depth feature learning, solves the problem of inconsistent data standards for different channels, and improves the feature expression ability. By adopting a multi-layer Transformer attention network and a TR-A2C deep reinforcement learning algorithm, the dynamic optimization and adaptive adjustment of risk control strategies are realized, breaking through the limitations of traditional static rule engines. Based on the stochastic network calculus method, the transaction processing delay is accurately modeled, and combined with a double-layer cache mechanism, the real-time response ability of the system in high-concurrency scenarios is ensured. Through the consistent hashing algorithm and the load balancing strategy, the distributed storage and dynamic scheduling of multi-channel data are realized, solving the problems of system scalability and data access performance. A complete data standardization processing process is designed, including protocol conversion, field mapping, and data verification, ensuring the consistency and reliability of multi-channel data. Through parallel policy optimization and a multi-environment sampling mechanism, the accuracy and real-time performance of risk control decisions are improved, effectively preventing transaction fraud risks. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the steps of an intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction in an embodiment of the present invention;
[0019] Figure 2 It is a structural block diagram of an intelligent bank payment and settlement device supporting real-time risk control and multi-channel interaction in an embodiment of the present invention;
[0020] Figure 3 It is a structural schematic diagram of a computer device in an embodiment of the present invention.
[0021] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments
[0022] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0023] Referring to Figure 1 , this embodiment provides an intelligent bank payment and settlement method supporting real-time risk control and multi-channel interaction, including the following steps:
[0024] S1. Perform multi-dimensional correlation analysis on multi-channel transaction behavior flow data to obtain transaction channel correlation data, and perform in-depth feature learning on the transaction channel correlation data to obtain transaction behavior feature vectors;
[0025] Among them, specific processing is performed on different dimensions of the multi-channel transaction behavior flow data to effectively extract valuable feature information. In the time dimension of the transaction data, the time slot division method is adopted, that is, the transaction data is sliced according to a preset time interval (such as minute level, hour level or day level), and the transaction is mapped to the corresponding time slot according to the time point when the transaction occurs, forming a binary feature matrix, where each time slot corresponds to a binary value, indicating whether a transaction has occurred during that time period, so as to effectively capture the distribution law of transactions in time. At the same time, in order to process the numerical information in the transaction data, such as transaction amount, transaction times or transaction frequency, one-hot encoding is used to discretize it, that is, according to the preset numerical interval, the continuous numerical data is divided into multiple intervals, and an independent binary feature is created for each interval to form a numerical feature distribution matrix, so as to avoid learning bias caused by uneven distribution of numerical data and enhance the model's perception ability of information such as transaction amount. When processing the geographical location information in the transaction data, the Euclidean distance between the transaction occurrence locations is calculated to measure the spatial distribution characteristics of the transaction behavior. In order to eliminate the influence between different geographical scales, the calculated Euclidean distance is normalized to obtain a position similarity matrix, which reflects the proximity of different transactions in the geographical space, thus helping to discover abnormal cross-regional transaction behaviors or potential fraud patterns. In order to improve the risk control ability, feature encoding is performed on the device identification information. The device fingerprint technology is adopted to extract stable features from information such as the device's IP address, browser fingerprint, operating system version, and screen resolution, and perform hash encoding or embedding mapping on them to obtain a device fingerprint feature matrix, which effectively represents the uniqueness of the device. Feature contribution degree calculation and data correlation analysis are performed on the binary feature matrix, numerical feature distribution matrix, position similarity matrix, and device fingerprint feature matrix to extract the key features of multi-channel transaction behavior. Feature contribution degree calculation uses methods based on information gain or Shapley value to measure the influence degree of various features on transaction behavior prediction, so as to screen out the most discriminative feature dimensions, while data correlation analysis uses methods such as association rule mining or graph neural networks to mine the internal relationships between different features to obtain transaction channel correlation data. The transaction channel correlation data is input into a multi-layer Transformer attention network for deep feature learning. The Transformer network dynamically weights the input data through the self-attention mechanism, so as to capture the long-range dependence relationship between transaction behavior features and avoid the defects of traditional sequence models in long-term dependence modeling. The multi-layer Transformer structure stacks multiple self-attention layers, and each layer performs multi-head attention calculation on the input data, so that the information between different transaction features interacts on different attention heads, and a non-linear transformation is performed through a feed-forward neural network to enhance the expression ability of the features.After learning through multiple Transformer networks, a transaction behavior feature vector is obtained, which is used for the training and inference of the risk control model to support real-time risk control decisions.
[0026] Input the transaction channel correlation data into the multi-head attention layer, and perform feature extraction through the multi-head attention mechanism. That is, in the multi-head attention layer, calculate the dot products of the query matrix, key matrix, and value matrix respectively through 8 parallel attention heads to obtain the attention weight distribution. Each attention head models the input data from different perspectives, capturing the dependencies of trading behaviors at different levels. Perform a scale transformation on the attention weight distribution to prevent the numerical values from being too large or too small and affecting the learning stability. Use a scaling factor to adjust the dot product results, and then apply softmax normalization to ensure that the sum of all attention weights is 1, obtaining a probability distribution, which enables the model to more reasonably allocate the weights of different trading features. Perform a weighted sum of the normalized attention weight distribution and the value matrix to obtain the multi-head attention features. Input the multi-head attention features into a two-layer feed-forward neural network to learn the deep features of trading behaviors. In the first layer of the feed-forward neural network, use the ReLU activation function to perform a non-linear transformation on the input data, enabling the model to learn the complex patterns of trading behaviors, enhancing the non-linear expression ability of the features, and improving the risk control system's ability to identify abnormal transactions. In the second layer of the feed-forward neural network, use a linear transformation to map the intermediate features to obtain the output features, which contain the deep representation of trading behaviors. To avoid the problems of gradient disappearance or gradient explosion and at the same time improve the training stability of the network, perform a residual connection on the output features. Add the input features and the output features after being transformed by the feed-forward neural network element by element, enabling the gradient to better retain the initial information during backpropagation and ensuring the training effect of the model. And perform layer normalization on the features after the residual connection to eliminate the distribution differences between different feature dimensions, making the training of the network more stable and enhancing the generalization ability of the model, obtaining the output feature vector of the first Transformer block. Input the output feature vector of the first Transformer block into the second Transformer block and the third Transformer block in sequence. In each Transformer block, repeat the execution of multi-head attention calculation, feed-forward neural network transformation, and layer normalization to ensure that each layer fully learns the deep features of trading behaviors. Through the design of the multi-layer structure, the model captures the multi-scale features of trading data at different levels and gradually learns the long-term dependencies of trading behaviors. Perform feature fusion and dimensionality reduction on the multi-layer feature representations to extract the most representative trading behavior feature vector. Feature fusion is carried out in the way of weighted sum or attention aggregation to integrate the feature information of different layers, making the final feature vector more comprehensively represent trading behaviors. Dimensionality reduction is carried out in the way of principal component analysis or full connection layer mapping to reduce the feature dimensions while retaining the most important feature information, and finally obtain the trading behavior feature vector.
[0027] S2. Perform protocol conversion and field mapping processing on multi-channel original message data to obtain standard payment transaction data;
[0028] Specifically, analyze the protocol types of multi-channel original message data and generate a protocol type identification set to determine the protocol category to which each message data belongs. By analyzing the protocol header, data structure, and field format, accurately identify the protocol type of the message and classify and store it in the protocol type identification set. According to the protocol type identification set, parse the multi-channel original message data in different protocol formats, split out the structured data fields, and generate an original field data set, which contains various transaction-related information, such as original fields like transaction amount, transaction time, merchant number, card number, user identity information, etc. Input the original field data set into the field mapping rule library for matching to determine how to convert the fields under different protocols into a standardized field structure. The field mapping rule library is a set of predefined mapping relation tables that store the mapping rules between the fields of different protocols and the standard payment data format. According to the standard field mapping table, perform field conversion on the original field data set to uniformly convert the field names and data formats under different protocols into the standard format to ensure that all payment transaction data follows the same field structure in subsequent processing. After this conversion step, obtain the initial standard data. Perform data integrity verification on the initial standard data to check for situations such as missing fields, incorrect data formats, or empty key fields. For example, the transaction amount field must be a positive number, the transaction time field needs to conform to the standard time format, and the merchant number must meet specific numbering rules, etc. Also perform consistency verification to check whether the data conforms to the business rules. Summarize the results of the integrity and consistency verification into the verification result data. Perform data completion and cleaning processing on the initial standard data according to the verification result data. If some fields are missing but can be completed by derivation, such as the missing transaction time being restored based on log records, then automatically fill in these data. Data cleaning is used to remove redundant or incorrect data, such as correcting the amount field with incorrect format, removing invalid characters, or standardizing the telephone number format, etc., and finally obtain a data format that conforms to the standard, forming the standard payment transaction data.
[0029] S3. Perform random network modeling on the multi-channel transaction processing flow based on the standard payment transaction data to obtain the processing delay threshold parameter;
[0030] It should be noted that the payment transaction processing flow is split into a data preprocessing module, a feature extraction module, a risk control decision-making module, and a response generation module. According to the functions and data flows of these modules, the processing node information in the standard payment transaction data is hierarchically divided to form a multi-level service node structure, so as to construct the topological relationship of the processing nodes. Through the hierarchical division of these service nodes, the transfer path of the payment transaction data in the entire system is described to ensure clear processing logics between different modules. Markov chain modeling is performed on the data transfer relationship between each node according to the processing node topological structure to construct the state transition matrix of the system. Markov chain modeling can describe the transfer probability of the payment transaction data between each processing node and quantify the dependence relationship between different nodes through the state transition matrix. Each processing node is regarded as a state, and the transfer probability of the payment transaction data between different states is determined by the state transition matrix. Thus, based on a large amount of historical transaction data, the possibility of different processing paths is estimated to generate a node dependence graph. This relationship graph reflects the logical connections between each processing node and provides the probability weights of different transaction flow paths. Maximum likelihood estimation and service time distribution parameter fitting are performed on the historical processing data of each node in the processing node topological structure to obtain the service time probability density function of each processing node, which describes the time distribution characteristics of the transaction data on different processing nodes. Laplace transform is performed on the node service time probability density function to calculate the delay cumulative effect of different nodes in the frequency domain space. The end-to-end delay distribution function is obtained by calculating the cumulative moment of the service time distribution, and this distribution function describes the completion probability of the entire transaction process at different time points. According to the preset service quality requirements, probability integral calculation is performed on the end-to-end delay distribution function to determine the delay constraint conditions that the transaction processing needs to meet. For example, if the system needs to ensure that 99% of the transactions can be completed within 200 milliseconds, the time threshold that meets this requirement is calculated based on the end-to-end delay distribution function. Numerical solution is performed according to the delay constraint conditions and the node dependence graph to calculate the upper bound of the optimal processing delay of the entire system. The numerical solution uses linear programming or convex optimization methods, with the goal of minimizing the maximum processing delay while satisfying all the delay constraint conditions. In this process, the optimal delay allocation strategy for transaction processing is obtained by optimizing the delay calculation of different paths, and finally the processing delay threshold parameter of the entire payment transaction process is calculated.
[0031] S4. Construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain the multi-channel real-time risk control decision-making strategy through parallel policy optimization.
[0032] Specifically, the transaction behavior feature vector is decomposed according to multiple dimensions such as transaction value, transaction frequency, transaction geographical location, and transaction device. Each dimension represents different types of transaction features. Among them, the transaction value reflects the change trend of the transaction amount, the transaction frequency describes the user's consumption habits, the transaction geographical location captures the spatial distribution pattern where the transaction occurs, and the transaction device provides device fingerprint features for identifying the stability and abnormal changes of the device. To consider the processing efficiency of transactions, a delay monitoring vector is constructed by combining the processing delay threshold parameter. This monitoring vector reflects the delay situation of each node in the transaction process, ensuring that the system effectively monitors transactions with excessive delays. The divided transaction feature subspaces and the delay monitoring vector are fused through the method of multi-dimensional tensor splicing to construct a state space embedding vector. Sparsity constraints and regularization processing are performed on the state space embedding vector to ensure that the model is not interfered by redundant features and focuses more on key transaction features. The sparsity constraint is achieved through L1 regularization to reduce the influence of irrelevant features, while the regularization processing adopts batch normalization or layer normalization methods to keep the features within a reasonable numerical range, thereby preventing gradient disappearance or gradient explosion. A risk control action space mapping is performed on the standardized state space, that is, an executable operation set of the risk control strategy is defined according to the state characteristics of the transaction. For example, directly approving the transaction, further authentication, triggering risk control rules, or rejecting the transaction, etc., to form an action space set, enabling the reinforcement learning model to make optimal decisions in different transaction states. When constructing the TR-A2C deep reinforcement learning network, the standardized state space is used as the input, and the network structure is designed to improve the learning efficiency. The Transformer encoding layer is used to deeply represent the state space, capturing the long-range dependencies between transaction features through the self-attention mechanism, and learning the interaction patterns between different features using the multi-head attention mechanism to generate high-quality state representations. The output of the Transformer encoding layer is connected to the policy layer and the value evaluation layer. The policy layer is used to generate the probability distribution of risk control decisions, while the value evaluation layer is used to calculate the long-term return of the state-action pair. By concatenating these three modules, the TR-A2C deep reinforcement learning network is formed, thus performing risk control decision training under the reinforcement learning framework. During the training process, parallel environment sampling is performed to improve the data collection efficiency and accelerate policy optimization. Multiple transaction simulation environments are created, and the current policy is run in each environment to collect interaction data. Each environment interacts according to the transaction state space and the action space, and generates an execution record of the transaction decision. These records are synchronously collected and stored as a policy trajectory sequence data set. This data set contains information such as transaction state, executed action, transaction processing time, and risk control feedback. The weighted combination of the delay loss and the risk control loss is calculated according to the policy trajectory sequence data set to optimize the reinforcement learning policy. Among them, the delay loss is used to measure the time cost of transaction processing, while the risk control loss is used to evaluate the security of the decision.To balance the accuracy of risk control and the efficiency of transaction processing, a weighted combination method is adopted to calculate the overall loss function, and the advantage function is calculated based on this loss to estimate the gain degree of each decision relative to the average decision, obtaining the policy gradient. The policy gradient is input into the Adam optimizer for parameter update, and the risk control decision-making ability of the model is continuously optimized through iterative loops until it converges to the optimal policy. Finally, a multi-channel real-time risk control decision-making strategy is obtained, enabling the system to intelligently identify abnormal transactions and make optimal risk control decisions in real-time transaction processing.
[0033] Perform access frequency statistics and data volume analysis on multi-channel transaction behavior flow data and multi-channel original message data to determine the access patterns and storage requirements of different types of data. By analyzing the access frequency of transaction behavior flow data, identify which data is frequently accessed in a short period of time, and which data belongs to historical transaction records that are stored for a long time but accessed less frequently. At the same time, conduct statistics on the access situation of the original message data to understand the distribution of transaction messages in different protocol formats during storage and calculation. Based on these analysis results, obtain the data classification results. Construct the data storage structure of the local cache space according to the data classification results to improve the access efficiency of frequently accessed data. Adopt the LRU algorithm for memory management and data elimination processing, that is, preferentially store the data with a higher recent access frequency in the local cache, and automatically eliminate the least recently accessed data when the cache space is insufficient, to ensure that frequently accessed data can always be retained in memory and improve the response speed of the system. Through this mechanism, establish a local cache data index table to enable the system to quickly locate and access cache data, reduce the latency of database queries, and improve the real-time performance of the payment transaction system. To optimize the distributed management of data storage, perform consistent hashing calculations on the storage nodes of the distributed cache layer to ensure the balanced distribution of data across multiple storage nodes. Adopt the consistent hashing algorithm to map M discrete points onto the hash ring and map virtual nodes to physical storage nodes to construct a node mapping relationship table. In this way, reduce the overhead of data migration and minimize the cost of data redistribution when storage nodes change (such as adding or removing nodes), thereby improving the scalability and stability of the system. After constructing the consistent hash ring, perform sharding processing on the historical transaction data to ensure the balance of data storage. Calculate the hash value according to the data sharding key (such as transaction ID, user ID, or timestamp), and determine the storage location of the data on the consistent hash ring to obtain the initial data distribution plan. Each piece of historical transaction data is automatically mapped to the corresponding storage node according to the hash calculation result, thus achieving the balanced storage of data and avoiding performance bottlenecks caused by excessive data concentration on specific storage nodes. In a distributed storage environment, to ensure that the computing resource utilization rate and network transmission latency of each storage node are in an optimal state, collect key performance indicators such as the computing resource occupancy rate and network latency of the storage nodes in real time, and calculate the health status of each node based on a weighted scoring model. This weighted scoring model combines multiple factors such as CPU usage, memory occupancy, disk I / O performance, and network bandwidth, assigns a health score to each storage node, constructs a load balancing scheduling table to reflect the current load situation of different storage nodes, so as to perform optimization adjustments when the storage node load is unbalanced. According to the load balancing scheduling table, perform dynamic migration of data shards to optimize the data storage distribution strategy.To reduce the overhead of data migration, the principle of minimizing migration cost is adopted. That is, when redistributing data, the amount of data movement is minimized as much as possible to reduce the system's computing and storage overhead. When performing data migration, storage nodes with higher health and lower load are preferentially selected to receive partial data shards, and the data shard storage scheme is updated after the migration is completed to ensure that the new data storage structure can adapt to the system's load conditions. Through this optimization strategy, efficient data storage and access capabilities are maintained in a dynamically changing transaction environment, thereby improving the computing efficiency of real-time risk control and the processing performance of multi-channel payment transactions.
[0034] In one example, multi-dimensional correlation analysis is performed on the multi-channel transaction behavior flow data to obtain transaction channel correlation data, and deep feature learning is performed on the transaction channel correlation data to obtain a transaction behavior feature vector, including:
[0035] The time dimension data in the multi-channel transaction behavior flow data is divided into time slots to obtain a binary feature matrix;
[0036] According to a preset numerical interval, one-hot encoding is performed on the numerical data in the multi-channel transaction behavior flow data to obtain a numerical feature distribution matrix;
[0037] Euclidean distance calculation and normalization are performed on the geographical location data in the multi-channel transaction behavior flow data to obtain a location similarity matrix;
[0038] Feature encoding is performed on the device identification information in the multi-channel transaction behavior flow data to obtain a device fingerprint feature matrix;
[0039] Feature contribution degree calculation and data correlation analysis are performed on the binary feature matrix, numerical feature distribution matrix, location similarity matrix, and device fingerprint feature matrix to obtain transaction channel correlation data;
[0040] The transaction channel correlation data is input into a multi-layer Transformer attention network for deep feature learning to obtain a transaction behavior feature vector.
[0041] In this example, the time dimension of the multi-channel transaction behavior flow data is processed to capture the time characteristics of the transaction behavior. The method of time slot division is adopted. According to the timestamp of the transaction occurrence, it is mapped to discrete time intervals, such as being divided at the hour, half-hour, or minute level to form a binary feature matrix. Suppose a certain transaction occurs at moment, and the entire time range is divided into time slots, then the time slot index corresponding to this transaction is expressed as:
[0042]
[0043] where, is the starting time of the time series, is the width of the time slot, and represents the time slot index where the transaction is located. Construct a binary vector , where only the position with index is set to 1, and the rest are set to 0, forming a binary feature matrix , where represents the transaction sample, represents the time slot. Process the numerical data. For continuous variables such as transaction amount and transaction frequency, use one-hot encoding to discretize them, map the numerical data according to the preset numerical intervals, and convert it into a numerical feature distribution matrix. Assume the transaction amount has a value range of , and it is divided into intervals, then the interval index where the transaction amount is located is expressed as:
[0044]
[0045] where, is the width of each interval. Construct a one-hot encoding vector , where the position with index is set to 1, and the rest are 0. In transaction behavior analysis, geographical location information is also a key feature, calculate the similarity between transaction locations. For the geographical location data and where the transaction occurs, use the Euclidean distance to measure the spatial distance between two points, and its calculation formula is:
[0046]
[0047] To ensure the balance between different data scales, normalize the Euclidean distance so that all distance values fall into the interval [0,1]. The normalization adopts the min-max scaling method, that is:
[0048]
[0049] where, and are respectively the minimum and maximum values among all calculated Euclidean distances. Construct a position similarity matrix , where each element Indicates the normalized geographical similarity. The closer the value is to 0, the closer the trading locations are. The closer the value is to 1, the farther the trading locations are. At the same time, feature encoding is performed on the device identification information to identify the device characteristics of the user and improve the accuracy of the risk control system. The information such as device ID, browser fingerprint, and operating system version is converted into a high-dimensional vector representation using hash encoding or embedding vector methods. For example, assume the device ID is , and through the hash function it is mapped to an embedding vector with a fixed dimension to construct a device fingerprint feature matrix . For the binary feature matrix , the numerical feature distribution matrix , the location similarity matrix , and the device fingerprint feature matrix , feature contribution degree calculation and data correlation analysis are carried out to construct transaction channel associated data. The feature contribution degree calculation uses methods such as information gain or Shapley value to measure the influence of each feature on the prediction of trading behavior. For example, assume the target variable of the trading behavior prediction model is , then the information gain calculation of a certain feature is:
[0050]
[0051] where is the entropy of the trading behavior, is the conditional entropy of the trading behavior given the feature . The greater the information gain, the greater the influence of the feature on the trading behavior. Based on the association rule mining method, the correlation between different features is analyzed to obtain transaction channel associated data. The transaction channel associated data is input into a multi-layer Transformer attention network for deep feature learning to obtain trading behavior feature vectors. The Transformer network calculates the importance of features through the self-attention mechanism and learns the dependency relationship between trading behavior features through the multi-head attention layer to calculate the attention weights:
[0052]
[0053] where are the query matrix, key matrix, and value matrix respectively, is the dimension of the key matrix. The multi-layer Transformer network extracts the global features of the trading behavior and generates the trading behavior feature vector .
[0054] In an example, the transaction channel associated data is input into a multi-layer Transformer attention network for deep feature learning to obtain trading behavior feature vectors, including:
[0055] Input the transaction channel association data into the multi - head attention layer, calculate the dot products of the query matrix, key matrix, and value matrix through 8 parallel attention heads to obtain the attention weight distribution;
[0056] Perform scale transformation and softmax normalization on the attention weight distribution, and perform weighted sum operation with the value matrix to obtain the multi - head attention features;
[0057] Input the multi - head attention features into a two - layer feed - forward neural network. The first layer uses the ReLU activation function for non - linear transformation to obtain intermediate features, and the second layer uses linear transformation to obtain output features;
[0058] Perform residual connection and layer normalization on the output features to obtain the output feature vector of the first Transformer block;
[0059] Input the output feature vector of the first Transformer block into the second Transformer block and the third Transformer block in sequence. Each layer of the Transformer block repeats the multi - head attention calculation, feed - forward network transformation, and normalization process to obtain multi - layer feature representations;
[0060] Perform feature fusion and dimensionality reduction on the multi - layer feature representations to obtain the transaction behavior feature vector.
[0061] In this example, the input transaction channel association data is structured to adapt to the calculation method of the multi - head attention mechanism. Let the transaction channel association data matrix be , where has a dimension of , where represents the number of transaction samples, represents the feature dimension of each transaction sample. To calculate the attention weights, is transformed into the query matrix , the key matrix and the value matrix , which is completed through linear transformation by three trainable weight matrices , and , and its calculation formula is:
[0062]
[0063] where, all have a dimension of , where is the dimension of a single attention head. To implement the multi - head attention mechanism, 8 independent attention heads are used, and each head has its own Perform a transformation so that the calculation is carried out in multiple subspaces to enhance the model's ability to extract features of trading behaviors. Calculate the query matrix and the key matrix to calculate their dot product and divide by the scaling factor to maintain numerical stability, obtaining the unnormalized attention weight distribution:
[0064]
[0065] where has a shape of and represents the importance of each trading sample to other samples. Apply softmax normalization to ensure that the sum of the attention weights is 1:
[0066]
[0067] where represents the normalized attention weight matrix. Multiply by the value matrix to perform a weighted sum to generate the multi-head attention features:
[0068]
[0069] where still has a shape of . For 8 attention heads, calculate respectively, and then concatenate them to form the multi-head attention output:
[0070]
[0071] where is a trainable weight matrix for mapping back to the original feature dimension and has a shape of . Input the multi-head attention features into a two-layer feed-forward neural network to extract deep features of trading behaviors. In the first layer of the feed-forward network, use the ReLU activation function for non-linear transformation:
[0072]
[0073] where has a dimension of , is the hidden layer dimension of the feed-forward network, is the bias term. In the second layer of the feed-forward network, adopt a linear transformation to obtain the final output features:
[0074]
[0075] Among them, has a dimension of , is the bias term, still maintains 's shape. To stabilize training and improve the performance of the model, residual connections and layer normalization are introduced during the calculation of the Transformer block. Residual connections ensure more stable gradient propagation between the input and output, and its calculation formula is:
[0076]
[0077] Apply layer normalization:
[0078]
[0079] Obtain the output feature vector of the first layer of the Transformer block, which contains the deep features of the trading behavior. Feed sequentially into the second and third layers of the Transformer block. In each layer, repeat the multi-head attention calculation, feed-forward network transformation, and normalization process to further extract the hierarchical features of the trading behavior. For the calculation of the second layer of the Transformer block:
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087] Similarly, the third layer of the Transformer block continues to process , and finally obtains the multi-layer feature representation of the complete trading behavior. To reduce the computational complexity and generate the trading behavior feature vector, feature fusion and dimensionality reduction are performed on the multi-layer feature representation. Use a fully connected layer for dimensionality reduction:
[0088]
[0089] Among them, has a dimension of , is the dimension of the final transaction behavior feature vector. As a transaction behavior feature vector, it is used for real-time risk control decision-making, enabling the payment system to efficiently identify abnormal transactions and improve the security and accuracy of payment transactions.
[0090] In one example, protocol conversion and field mapping processing are performed on multi-channel raw message data to obtain standard payment transaction data, including:
[0091] Perform protocol type identification on multi-channel raw message data to obtain a protocol type identification set, and parse and split the multi-channel raw message data according to the protocol type identification set to obtain a raw field data set;
[0092] Input the raw field data set into the field mapping rule library for matching to obtain a standard field mapping table, and perform field conversion on the raw field data set according to the standard field mapping table to obtain initial standard data;
[0093] Perform data integrity verification and consistency check on the initial standard data to obtain verification result data, and perform data completion and cleaning processing on the initial standard data according to the verification result data to obtain standard payment transaction data.
[0094] In this example, protocol type analysis is performed on multi-channel raw message data to determine the communication protocol formats used for transaction data from different sources. Let the raw message data set be , where represents the raw message of the th transaction, and transaction data from different channels uses different protocol formats, such as ISO 8583, XML, JSON, or fixed-length messages. Build a protocol identification model to automatically determine the protocol type to which a message belongs based on its structure and field characteristics:
[0095]
[0096] where, represents the protocol type of the th transaction, to obtain a protocol type identification set , which is used to guide the subsequent parsing process. According to the protocol type identification set, the multi-channel raw message data is parsed and split to extract transaction-related field information. For structured protocols (such as JSON, XML), key-value pair fields are extracted based on a hierarchical parser:
[0097]
[0098] where, represents the A set of fields for a transaction. For fixed - length formats (such as ISO 8583 messages), field splitting is based on predefined offsets:
[0099]
[0100] Among them, and represent the start and end offsets of the th field respectively, to obtain the complete original field data set . Perform field standardization processing on the original field data to ensure that fields in different protocol formats are mapped to a unified data structure. Build a field mapping rule library , which stores the field correspondence relationships under different protocol formats. For example:
[0101]
[0102] Input the original field data set into the field mapping rule library for matching to generate a standard field mapping table:
[0103]
[0104] Among them, represents the standard field mapping table for the th transaction. Based on this mapping table, perform field conversion on the original field data set to obtain the initial standard data:
[0105]
[0106] Among them, represents the standardized transaction data after conversion. Perform data integrity and consistency verification on the standardized transaction data. The purpose of integrity verification is to check whether there are missing key fields. For the th transaction, define an integrity verification function:
[0107]
[0108] Among them, is the index set of all key fields, is an indicator function. When a certain key field is empty, , it indicates that the integrity verification fails. The goal of consistency verification is to check whether the field values conform to the expected rules. Let the consistency verification function be:
[0109]
[0110] Among them, represents the format verification rule for a specific field. If all fields meet the format requirements, then otherwise The comprehensive result of integrity and consistency verification is:
[0111]
[0112] When it indicates that the data passes the verification, otherwise data completion and cleaning are required. For the data that fails the verification, historical data or business rules are used for completion. For example, if the transaction time is missing, the time information embedded in the transaction serial number is used for filling. Abnormal data is cleaned, such as removing non-standard characters, normalizing currency units, and after cleaning, standard payment transaction data is obtained.
[0113] In one example, a random network model is built for the multi-channel transaction processing flow based on the standard payment transaction data to obtain the processing delay threshold parameters, including:
[0114] Based on the data preprocessing module, feature extraction module, risk control decision-making module, and response generation module, the processing node information in the standard payment transaction data is hierarchically divided and multi-level service nodes are constructed to obtain the processing node topology structure;
[0115] Based on the processing node topology structure, a Markov chain model is built for the data flow relationship between nodes and a state transition matrix is constructed to obtain the node dependency graph;
[0116] Perform maximum likelihood estimation and service time distribution parameter fitting on the historical processing data of each node in the processing node topology structure to obtain the node service time probability density function;
[0117] Perform Laplace transform on the node service time probability density function and calculate the cumulative moment of the service time distribution to obtain the end-to-end delay distribution function, and perform probability integral calculation on the end-to-end delay distribution function according to the preset service quality requirements to obtain the delay constraint condition;
[0118] Based on the delay constraint condition and the node dependency graph, perform numerical solution and calculate the upper bound of the optimal processing delay to obtain the processing delay threshold parameter.
[0119] In this example, a structured model is built for the processing flow of payment transaction data to analyze the execution logic and data flow mode of each processing node. Let the standard payment transaction data set be , where each Represents the complete data of a transaction, including user identity information, transaction amount, transaction time, payment channel, device information, etc. During the actual operation of the payment system, each transaction goes through multiple processing links, including data preprocessing, feature extraction, risk control decision-making, and finally response generation. These modules are used as the main processing layers, and their internal nodes are refined to construct a multi-level service node structure. To establish the processing node topology, assume that the processing system contains different sets of processing nodes , where each node corresponds to a specific computing task, such as data parsing, risk assessment, transaction confirmation, etc. To construct the connection relationship between nodes, define the data transfer function to represent the data transfer relationship from node to node , that is:
[0120]
[0121] Through this step, the processing flow of the entire payment transaction system is represented as a directed graph , where is the set of all processing nodes, and the edge set is composed of all node pairs that satisfy , forming a complete processing node topology. Based on the processing node topology, model the data transfer relationship between each node to analyze the flow characteristics of transactions between different processing nodes. Use Markov chains to describe the state transition process of data transfer. Assume that the state set in the system is , and define the state transition matrix :
[0122]
[0123] Among them, represents the probability that the transaction data flows from node to , which is estimated through historical data. Assume that in the historical transaction records, there are a total of times that the data stream is transmitted from node to , then the state transition probability is calculated as:
[0124]
[0125] Finally, obtain the state transition matrix , used to describe the data flow pattern in the payment system and construct a node dependency graph. Model the service time distribution of each node in the processing node topology to calculate the delay characteristics of the system. Assume that each node The time to process a transaction follows a probability distribution , and use the maximum likelihood estimation method to fit its parameters. Let be the node 's historical processing time samples, then the goal of maximum likelihood estimation is to find the best distribution parameters such that:
[0126]
[0127] Common distribution fitting methods include exponential distribution, normal distribution or gamma distribution. For example, if the processing time of a node follows an exponential distribution:
[0128]
[0129] where, represents the probability density at the node ; is the processing rate parameter, indicating the probability that the node completes a transaction per unit time; the exponential distribution is characterized by memorylessness, that is, the current transaction processing time does not affect the future transaction processing time. Estimate , using the "maximum likelihood estimation" method. Let the set of historical processing time samples of the node be:
[0130]
[0131] where, represents the th transaction's processing time at the node . Then the optimal parameter of the exponential distribution is calculated as:
[0132]
[0133] If historical data shows that the processing time better fits a normal distribution:
[0134]
[0135] where, represents the average processing time, that is, the average duration of the node to process a transaction; represents the variance of the processing time, characterizing the degree of fluctuation of the node's processing time. Then the parameter estimation adopts:
[0136]
[0137] Obtain the probability density function of the service time for all nodes. After obtaining the service time distribution of the nodes, calculate the end-to-end delay distribution of the entire system. Perform the Laplace transform on the probability density function of the service time for each node:
[0138]
[0139] Then the Laplace transform of the end-to-end delay is expressed as:
[0140]
[0141] For Perform the inverse Laplace transform to obtain the delay distribution function of the trading system and calculate the cumulative distribution function:
[0142]
[0143] where represents the maximum allowable delay; if meets the quality of service requirements (e.g., 99% of the transactions are completed within ), then the system performance is acceptable, otherwise optimization is required. According to the delay constraint conditions and the node dependency graph, calculate the optimal upper bound of the processing delay. Let the optimal delay allocation for each node be , then the optimization objective is expressed as:
[0144]
[0145] subject to
[0146]
[0147] After solving this optimization problem, obtain the processing delay threshold parameter:
[0148]
[0149] This value represents the maximum transaction processing time of the entire system in the optimal state and is used to dynamically adjust the computing resources to optimize the response speed of the payment system.
[0150] In an example, construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain the multi-channel real-time risk control decision-making strategy through parallel policy optimization, including:
[0151] The transaction behavior feature vector is divided into feature subspaces according to transaction value, transaction frequency, transaction geographical location, and transaction device, and a delay monitoring vector is constructed based on the processing delay threshold parameter. A state space embedding vector is obtained through multi-dimensional tensor splicing;
[0152] Perform sparsity constraint and regularization processing on the state space embedding vector to obtain a standardized state space, and map the standardized state space to a risk control action space to obtain a set of action spaces;
[0153] Based on the standardized state space, connect the Transformer encoding layer, policy layer, and value evaluation layer in series to obtain a TR-A2C deep reinforcement learning network, and perform parallel environment sampling on the TR-A2C deep reinforcement learning network. A policy trajectory sequence dataset is obtained by synchronously collecting interaction data from multiple environments;
[0154] Calculate the weighted combination of delay loss and risk control loss based on the policy trajectory sequence dataset, and calculate the advantage function based on the weighted combination to obtain the policy gradient;
[0155] Input the policy gradient into the Adam optimizer for parameter update and iterative loop to obtain a multi-channel real-time risk control decision-making strategy.
[0156] In this example, a reasonable feature subspace division is performed on the transaction behavior feature vector so that the model can learn different dimensional information of transaction data and make more accurate risk control decisions during the reinforcement learning process. Let the transaction behavior feature vector be:
[0157]
[0158] Among them, represents the feature vector of the th transaction, and is the total number of transaction samples. To make full use of the information of transaction data, it is divided according to different feature categories of transactions to obtain four feature subspaces: the transaction value feature subspace , which contains numerical information such as transaction amount and payment method; the transaction frequency feature subspace , which contains the number of recent transactions of the user, average transaction interval, etc.; the transaction geographical location feature subspace , which contains the longitude and latitude where the transaction occurs, transaction distance, etc.; the transaction device feature subspace , which contains device fingerprint, operating system information, etc. Let represent the dimensions of each feature subspace respectively, then it is expressed as:
[0159]
[0160] Meanwhile, to measure the processing latency of transactions, a latency monitoring vector is constructed based on the processing latency threshold parameter , where each element of represents the processing latency of the th transaction:
[0161]
[0162] Among them, is obtained by normalizing the actual processing latency of the transaction:
[0163]
[0164] Among them, and represent the minimum and maximum latencies observed in the transaction system respectively. The above-mentioned feature subspace is concatenated with the latency monitoring vector to construct a state space embedding vector:
[0165]
[0166] Among them, serves as the input state vector for reinforcement learning. Sparse constraints and regularization processing are performed on the state space embedding vector. Define the L1 regularization term:
[0167]
[0168] Among them, is the regularization coefficient, represents the L1 norm of the state vector, which is used to encourage the sparsity of features and make the weights of irrelevant information tend to zero. Perform the normalization operation:
[0169]
[0170] Among them, and are the mean and standard deviation of respectively. After the above processing, the standardized state space is obtained and mapped to the risk control action space, that is, the risk control decision set is defined:
[0171]
[0172] Among them, represents the number of risk control strategies. Based on the standardized state space Design the TR-A2C deep reinforcement learning network, which consists of an encoding layer, a policy layer, and a value evaluation layer. The encoding layer uses the self-attention mechanism to process the transaction state information:
[0173]
[0174] Among them, are the query matrix, the key matrix, and the value matrix respectively, is the dimension of the attention head. The policy layer generates the probability distribution of the risk control decision:
[0175]
[0176] Among them, is the neural network used to calculate the risk control policy score. At the same time, the value evaluation layer calculates the state value function:
[0177]
[0178] To train the TR-A2C network, parallel environment sampling is performed to collect reinforcement learning trajectories simultaneously in multiple transaction environments:
[0179]
[0180] Among them, is the reward for the transaction decision. Define the latency loss and the risk control loss :
[0181]
[0182]
[0183] Among them, is the target latency threshold, is the current transaction processing latency. Calculate the weighted combined loss:
[0184]
[0185] Among them, and are the weight hyperparameters. Calculate the advantage function to update the policy:
[0186]
[0187] Among them, is the advantage function, measuring the advantage of the current action ; is the discount factor, used to measure the impact of future rewards on the current decision, and the value range is ; is the expected value of the next state; is the expected value of the current state; is the immediate reward, representing the risk control benefit of the current transaction. If , it indicates that the currently selected risk control strategy is better than the average strategy; if , it means that the benefit of the current action is lower than the average value, and the strategy needs to be adjusted. Calculate the policy gradient:
[0188]
[0189] where is the policy objective function, representing the expected return of the policy ; are the parameters of the policy network, i.e., the neural network weights used to optimize risk control decisions; is the gradient with respect to the parameter , used to update the policy network to increase the probability of high-return actions; represents the expectation over all transaction samples; is the action probability distribution of the current policy network, indicating the probability of selecting in state , calculate the log probability of this action; is the advantage function, used to measure the improvement direction of the current policy. Execute parameter update using an optimizer:
[0190]
[0191] where are the parameters of the policy network; is the learning rate, controlling the step size of gradient update; is the policy gradient, used to optimize the policy network to make it more biased towards high-benefit decisions. After multiple rounds of iterative loops, the optimal multi-channel real-time risk control decision-making strategy is finally obtained, automatically selecting risk control actions according to the transaction state, realizing intelligent interception of fraudulent transactions, and improving the security and transaction experience of the payment system.
[0192] Among them, the policy gradient is input into the Adam optimizer for parameter update and iterative loop to obtain a multi-channel real-time risk control decision-making strategy, including: constructing a time-varying Riccati equation for the convergence rate of the policy gradient and the decision amplitude constraint, and designing a time-varying parameter scheduler according to the system state error and the decision output boundary to obtain a parameter update matrix; performing online parameter adjustment on the policy layer and the value evaluation layer of the TR-A2C deep reinforcement learning network based on the parameter update matrix to obtain a time-varying gain sequence; inputting the time-varying gain sequence into the momentum term and the second-moment estimation term of the Adam optimizer to calculate the adaptive learning rate and obtain an optimizer parameter group; performing normalization processing and amplitude limit on the policy gradient according to the optimizer parameter group, and obtaining a constrained gradient vector through non-linear projection; sampling the constrained gradient vector according to the preset batch size and calculating the sample weights to obtain a weighted training data set; performing mini-batch stochastic gradient descent on the weighted training data set and estimating the convergence time through linear regression to obtain a time-varying learning rate sequence; dynamically adjusting the update step size of the optimizer according to the time-varying learning rate sequence and setting an early stopping condition for convergence judgment to obtain an optimal parameter solution; constructing a risk control decision function based on the optimal parameter solution and determining the convergence boundary through finite-time stability analysis to obtain a multi-channel real-time risk control decision-making strategy.
[0193] In one example, the intelligent bank payment and settlement method that supports real-time risk control and multi-channel interaction further includes:
[0194] Performing access frequency statistics and data volume analysis on multi-channel transaction behavior flow data and multi-channel original message data to obtain a data classification result;
[0195] Constructing a data storage structure for the local cache space according to the data classification result, and performing memory management and data elimination processing through the LRU algorithm to obtain a local cache data index table;
[0196] Performing consistent hashing calculation on the storage nodes of the distributed cache layer, generating a hash ring composed of M discrete points, and mapping virtual nodes to physical nodes to obtain a node mapping relationship table;
[0197] Calculating the hash value of the historical transaction data according to the data sharding key, and determining the data storage location on the consistent hash ring to obtain an initial data distribution scheme;
[0198] Real-time collecting the computing resource occupancy rate and network transmission delay of each storage node, and calculating the node health degree based on a weighted scoring model to obtain a load balancing scheduling table;
[0199] Performing data sharding dynamic migration according to the load balancing scheduling table, and reallocating the data storage location according to the principle of minimum migration cost to obtain a multi-channel data sharding storage scheme.
[0200] In this example, the access pattern of the data is analyzed to formulate an optimal storage and caching strategy. Let the transaction behavior flow data set be:
[0201]
[0202] Among them, represents the th complete record of a transaction, including information such as transaction time, transaction amount, transaction device, transaction channel, etc. For the original message data set, let it be:
[0203]
[0204] Among them, represents the original message from different payment channels. The access frequency and data volume of the above data are counted. Let the access frequency function be:
[0205]
[0206] Among them, represents the number of times the transaction is accessed within the time window . Similarly, for data volume analysis, a data size function is defined:
[0207]
[0208] Among them, represents the th field of the transaction, is the total number of fields, is the field size. By calculating the and of all transactions, data grading is performed. According to the data grading results, a local cache structure is built for high-frequency data to improve query efficiency and reduce database pressure. Let the cache space size be , then the cache data set is:
[0209]
[0210] Among them, is the access frequency of the lowest cache requirement. To optimize cache utilization, the LRU algorithm is used for memory management, that is, when the cache space reaches the upper limit, the least recently used data is removed. Let the most recent access time of the cache data be , then the LRU eviction policy is defined as:
[0211]
[0212] That is, the earliest unaccessed data is removed, and in the cache data index table Maintain the latest cache records:
[0213]
[0214] After optimizing the local cache, optimize the distributed storage layer, perform consistent hashing calculation, and generate a hash ring to ensure that data is evenly distributed on storage nodes. Let the set of storage nodes be:
[0215]
[0216] Through the hash function Calculate the position of each node:
[0217]
[0218] And map hash values to the hash ring to minimize the data migration cost. Introduce virtual nodes where each physical node is mapped to multiple virtual nodes:
[0219]
[0220] Thus effectively reducing the data skew problem and forming a node mapping relationship table:
[0221]
[0222] Calculate the hash value according to the data sharding key to determine the storage location of the transaction data. Let the set of key fields of the transaction data be , define the data sharding key:
[0223]
[0224] Then find the nearest storage node on the hash ring:
[0225]
[0226] Ensure that each transaction data is evenly distributed to the appropriate storage nodes and form an initial data distribution scheme. During the data storage operation, monitor the computing resource occupancy rate and network transmission latency of the storage nodes in real time to evaluate the node health. Let the node have a CPU usage rate of , a memory usage rate of , a disk I / O rate of , and a network latency of , then the node health calculation is:
[0227]
[0228] Among them, is a weight parameter, reflecting the impact of different resources on the system performance. After calculating the health of all storage nodes, a load balancing scheduling table is generated:
[0229]
[0230] And the data storage is readjusted according to the load balancing policy. In order to reduce the data migration cost, the principle of minimum migration cost is adopted for dynamic data shard migration. Suppose the current storage node has a load exceeding the threshold, then part of the data is migrated to a node with higher health. Define the migration cost:
[0231]
[0232] Among them, represents the amount of migrated data; represents the additional delay brought by migration; and are weight parameters. The optimal migration target node is determined by:
[0233]
[0234] Through this step, the data will be migrated to the target node with the minimum cost, and finally a multi-channel data shard storage scheme is formed.
[0235] Referring to Figure 2 , this embodiment provides an intelligent bank payment and settlement device supporting real-time risk control and multi-channel interaction, including:
[0236] The association analysis unit 1 is used to perform multi-dimensional association analysis on the multi-channel transaction behavior flow data to obtain transaction channel association data, and perform in-depth feature learning on the transaction channel association data to obtain transaction behavior feature vectors;
[0237] The conversion unit 2 is used to perform protocol conversion and field mapping processing on the multi-channel original message data to obtain standard payment transaction data;
[0238] The modeling analysis unit 3 is used to perform stochastic network modeling on the multi-channel transaction processing flow according to the standard payment transaction data to obtain the processing delay threshold parameter;
[0239] The risk control decision-making unit 4 is used to construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain the multi-channel real-time risk control decision-making strategy through parallel policy optimization.
[0240] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment and will not be elaborated herein.
[0241] Refer to Figure 3 , in the embodiment of the present invention, a computer device is further provided. The computer device may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0242] Those skilled in the art can understand that Figure 3 the structure shown in
[0243] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0244] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0245] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method including that element.
[0246] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. An intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction, characterized in that: The following steps are involved: Performing multi-dimensional correlation analysis on multi-channel transaction behavior flow data to obtain transaction channel correlation data, and performing deep feature learning on the transaction channel correlation data to obtain a transaction behavior feature vector; Perform protocol conversion and field mapping on original message data from multiple channels to obtain standard payment transaction data; Performing random network modeling on the multi-channel transaction processing flow according to the standard payment transaction data to obtain a processing delay threshold parameter; The state space and reward function of the TR-A2C deep reinforcement learning network are constructed according to the transaction behavior feature vector and the processing delay threshold parameter, and a multi-channel real-time risk control decision strategy is obtained through parallel strategy optimization; specifically, the transaction behavior feature vector is divided into feature subspaces according to transaction value, transaction frequency, transaction geographic location and transaction device, and a delay monitoring vector is constructed based on the processing delay threshold parameter, and a state space embedding vector is obtained by multi-dimensional tensor splicing; sparsity constraints and regularization are performed on the state space embedding vector to obtain a standardized state space, and the standardized state space is mapped to the risk control action space , and obtain an action space set; based on the standardized state space, the Transformer encoding layer, the policy layer and the value assessment layer are connected in series to obtain a TR-A2C deep reinforcement learning network, and the TR-A2C deep reinforcement learning network is sampled in parallel environments, and interaction data is synchronously collected through multiple environments to obtain a strategy trajectory sequence data set; according to the strategy trajectory sequence data set, a weighted combination of delay loss and risk control loss is calculated, and an advantage function is calculated based on the weighted combination to obtain a policy gradient; the policy gradient is input into the Adam optimizer for parameter update and cyclic iteration to obtain a multi-channel real-time risk control decision strategy.
2. The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction according to claim 1, characterized in that: The multi-dimensional correlation analysis is performed on the multi-channel transaction behavior flow data to obtain transaction channel correlation data, and deep feature learning is performed on the transaction channel correlation data to obtain a transaction behavior feature vector, including: Divide the time dimension data in the multi-channel transaction behavior flow data into time slots to obtain a binary feature matrix; According to the preset numerical range, the numerical data in the multi-channel transaction behavior flow data is uniquely encoded to obtain the numerical feature distribution matrix; The Euclidean distance calculation and normalization are performed on the geographical location data in the multi-channel transaction behavior flow data to obtain the location similarity matrix; Perform feature encoding on the device identification information in the multi-channel transaction behavior flow data to obtain the device fingerprint feature matrix; Performing feature contribution calculation and data association analysis on the binary feature matrix, the numerical feature distribution matrix, the position similarity matrix, and the device fingerprint feature matrix to obtain transaction channel association data; The transaction channel association data is input into a multi-layer Transformer attention network for deep feature learning to obtain a transaction behavior feature vector.
3. The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction according to claim 2 is characterized in that: The step of inputting the transaction channel association data into a multi-layer Transformer attention network for deep feature learning to obtain a transaction behavior feature vector includes: Input the transaction channel association data into the multi-head attention layer, and calculate the dot product of the query matrix, the key matrix and the value matrix through 8 parallel attention heads to obtain the attention weight distribution; The attention weight distribution is scaled and softmax normalized, and a weighted sum operation is performed with the value matrix to obtain a multi-head attention feature; The multi-head attention features are input into a two-layer feedforward neural network, the first layer uses a ReLU activation function to perform nonlinear transformation to obtain intermediate features, and the second layer uses a linear transformation to obtain output features; Performing residual connection and layer normalization processing on the output features to obtain an output feature vector of the first layer Transformer block; The output feature vector of the first-layer Transformer block is sequentially input into the second-layer Transformer block and the third-layer Transformer block. Each layer of Transformer block repeatedly performs multi-head attention calculation, feedforward network transformation and normalization processing to obtain a multi-layer feature representation; The multi-layer feature representation is subjected to feature fusion and dimensionality reduction processing to obtain a transaction behavior feature vector.
4. The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction according to claim 1, characterized in that: The protocol conversion and field mapping processing of the original message data from multiple channels to obtain standard payment transaction data includes: Performing protocol type identification on the original message data of multiple channels to obtain a protocol type identification set, and parsing and splitting the original message data of multiple channels according to the protocol type identification set to obtain an original field data set; Input the original field data set into the field mapping rule library for matching to obtain a standard field mapping table, and perform field conversion on the original field data set according to the standard field mapping table to obtain initial standard data; The initial standard data is subjected to data integrity check and consistency check to obtain verification result data, and the initial standard data is subjected to data completion and cleaning processing according to the verification result data to obtain standard payment transaction data.
5. The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction according to claim 1, characterized in that: The performing random network modeling on the multi-channel transaction processing flow according to the standard payment transaction data to obtain a processing delay threshold parameter includes: According to the data preprocessing module, the feature extraction module, the risk control decision module and the response generation module, the processing node information in the standard payment transaction data is hierarchically divided and a multi-level service node is constructed to obtain a processing node topology structure; According to the processing node topology structure, Markov chain modeling and state transfer matrix construction are performed on the data flow relationship between nodes to obtain a node dependency graph; Performing maximum likelihood estimation and service time distribution parameter fitting on the historical processing data of each node in the processing node topology structure to obtain a node service time probability density function; Performing Laplace transformation on the node service time probability density function and calculating the cumulative moment of the service time distribution to obtain an end-to-end delay distribution function, and performing probability integral calculation on the end-to-end delay distribution function according to a preset service quality requirement to obtain a delay constraint condition; Numerical solution and upper bound calculation of optimal processing delay are performed according to the delay constraint condition and the node dependency graph to obtain a processing delay threshold parameter.
6. The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction according to claim 1, characterized in that: The intelligent bank payment settlement method supporting real-time risk control and multi-channel interaction also includes: Performing access frequency statistics and data volume analysis on the multi-channel transaction behavior flow data and the multi-channel original message data to obtain data classification results; The data storage structure of the local cache space is constructed according to the data classification result, and the memory management and data elimination processing are performed through the LRU algorithm to obtain the local cache data index table; Perform consistent hash calculation on the storage nodes of the distributed cache layer to generate M discrete points to form a hash ring, and map the virtual nodes to the physical nodes to obtain a node mapping relationship table; Calculate the hash value of historical transaction data according to the data sharding key, determine the data storage location on the consistent hash ring, and obtain the initial data distribution plan; The computing resource occupancy rate and network transmission delay of each storage node are collected in real time, and the node health is calculated based on the weighted scoring model to obtain the load balancing schedule; Perform dynamic migration of data shards according to the load balancing schedule, reallocate data storage locations according to the principle of minimum migration cost, and obtain a multi-channel data shard storage solution.
7. An intelligent bank payment and settlement device supporting real-time risk control and multi-channel interaction, characterized in that: For implementing the steps of the method according to any one of claims 1 to 6, the device comprises: A correlation analysis unit, used to perform multi-dimensional correlation analysis on multi-channel transaction behavior flow data to obtain transaction channel correlation data, and perform deep feature learning on the transaction channel correlation data to obtain a transaction behavior feature vector; A conversion unit, used to perform protocol conversion and field mapping processing on multi-channel original message data to obtain standard payment transaction data; A modeling and analysis unit, configured to perform random network modeling on a multi-channel transaction processing flow according to the standard payment transaction data to obtain a processing delay threshold parameter; The risk control decision unit is used to construct the state space and reward function of the TR-A2C deep reinforcement learning network according to the transaction behavior feature vector and the processing delay threshold parameter, and obtain a multi-channel real-time risk control decision strategy through parallel strategy optimization.
8. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and system for realizing electronic channel risk control disposal based on code insertion technology
CN119538269A