Cross-border B2B2C account division and settlement system based on reinforcement learning

The cross-border B2B2C revenue sharing and settlement system, which utilizes reinforcement learning and blockchain technology, solves the complexity and compliance issues of revenue sharing and settlement in the cross-border B2B2C business model, and achieves efficient and accurate fund management and a secure settlement process.

CN120875866AInactive Publication Date: 2025-10-31LIANDUODUO INFORMATION TECH (CHONGQING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511365952.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The revenue sharing and settlement processes in the cross-border B2B2C business model are complex. Manual processing leads to high error rates, long settlement cycles, difficulty in tracing fund flows, and significant compliance risks.

Method used

A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning is adopted, integrating smart contracts and automated calculation modules to achieve flexible configuration of multi-level revenue sharing rules and fully automated revenue sharing. It combines blockchain technology for tamper-proof evidence storage, supports currency conversion and multi-country tax calculation, and optimizes revenue sharing rules using an improved bidirectional long short-term memory network and a deep deterministic policy gradient algorithm.

Benefits of technology

It significantly improves the processing efficiency and accuracy of cross-border B2B2C platforms, shortens settlement cycles, reduces capital occupation costs, enhances system security and compliance, supports multi-currency settlement and customized revenue sharing schemes, and adapts to diverse transaction needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875866A_ABST
    Figure CN120875866A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based cross-border B2B2C accounting and settlement system, and the system comprises the following modules: a transaction information collection module which achieves the collection of cross-border order information and the generation of an order data set; the exchange rate conversion module is used for realizing external exchange rate data acquisition and settlement currency conversion of the order amount; the tax calculation module is responsible for being connected with a tax policy interface and calculating order tax; the intelligent contract module is used for realizing structured analysis of account division rules and intelligent optimization of account division parameters; the order risk judgment module is used for realizing automatic evaluation of order states and risk characteristics and account division instruction triggering; the separate account calculation module is used for completing automatic calculation and result verification of separate account amounts; the settlement module is used for automatically completing multi-party settlement and account arrival state recording of sub-account funds; and the data storage and tracing module is used for realizing whole-process data evidence storage and traceable management. According to the cross-border B2B2C transaction system and method, multi-level account division automation is realized, and the efficiency, the safety and the compliance of cross-border B2B2C transactions are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cross-border e-commerce financial technology, and in particular to a cross-border B2B2C revenue sharing and settlement system based on reinforcement learning. Background Technology

[0002] In cross-border B2B2C business models, transactions often involve multiple levels of entities, including platforms, suppliers, and distributors. The revenue sharing and settlement process is a common pain point in the industry due to the large number of participants and the complexity of revenue sharing rules. Revenue sharing agreements between different entities vary significantly, with revenue sharing ratios fluctuating between 5% and 30%. Distributor commissions are often linked to sales performance and require dynamic adjustments. With manual methods, handling over a thousand transactions daily, matching and processing rules one by one is prone to errors, with error rates as high as 3% to 5%, leading to large-scale transaction disputes, lengthy error correction cycles, and severely impacting business efficiency.

[0003] Cross-border settlement also involves complex processes such as currency exchange, calculation of taxes in multiple countries, and cross-border bank transfers. Under traditional manual processing methods, currency conversion and tax matching require manual operation, resulting in settlement cycles of 7-15 days. This directly slows down suppliers' cash flow and significantly increases companies' capital costs. Furthermore, existing systems lack complete transaction data records and tamper-proof storage mechanisms, making the accounting and settlement process susceptible to human intervention. The flow of funds is difficult to trace, and once abnormal fund operations occur, it is difficult to hold those responsible accountable in a timely manner, highlighting significant compliance risks.

[0004] Therefore, how to provide a cross-border B2B2C revenue sharing and settlement system based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a reinforcement learning-based cross-border B2B2C revenue sharing and settlement system. By integrating smart contracts and automated calculation modules, this invention achieves flexible configuration of multi-level revenue sharing rules and fully automated revenue sharing throughout the entire process, significantly improving the processing efficiency and accuracy of cross-border B2B2C platforms. It also integrates currency conversion and multi-country tax calculation functions, significantly shortening the settlement cycle. Utilizing blockchain technology for immutable storage of revenue sharing and settlement data ensures full traceability and regulatory compliance throughout the revenue sharing process, further guaranteeing fund security. The system also supports multi-currency settlement, customized revenue sharing schemes, and adaptation to multiple tax systems, flexibly meeting diverse cross-border business scenarios and varied transaction needs, comprehensively enhancing the intelligence and standardization of supply chain finance and fund management.

[0006] A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to an embodiment of the present invention includes the following modules: The transaction information collection module is used to acquire cross-border order information and generate order datasets; The currency conversion module is used to obtain exchange rate data in real time through the interface of external financial institutions, and convert the order amount into the settlement currency amount according to the exchange rate to obtain the currency conversion result; The tax calculation module is used to connect to external tax policy interfaces, calculate the comprehensive taxes payable based on the order dataset and exchange rate conversion results, and obtain the tax amount. The smart contract module is used to call the natural language rule conversion algorithm to parse the input revenue sharing rule text into a structured executable rule expression. The natural language rule conversion algorithm adopts an improved bidirectional long short-term memory network with dispute memory gating, extracts features from the positive and negative context information respectively, obtains the positive and negative hidden state sequences, and combines historical transaction data to dynamically optimize the revenue sharing rule parameters through a deep deterministic strategy gradient algorithm to generate the revenue sharing rule parameters applicable to the current order. The order risk assessment module is used to determine the order status and risk characteristics. It adopts a gradient boosting tree model to determine the order status and risk characteristics. If the revenue sharing conditions are met, a revenue sharing calculation instruction is generated. The revenue sharing calculation module is used to respond to revenue sharing calculation instructions, calculate the revenue sharing amount of each entity based on the exchange rate conversion result, tax amount and revenue sharing rules, and verify the sum of the revenue sharing amounts of each entity. If the verification passes, a settlement execution instruction is generated. The settlement module is used to connect to the interface of external payment institutions, transfer the split amount to the corresponding entity account, and record the arrival time and settlement status; The data storage and traceability module is used to store order datasets, currency conversion results, tax amounts, revenue sharing rules, revenue sharing amounts, settlement results, and related logs.

[0007] Furthermore, the modules are interconnected using the following method: The system obtains cross-border order information, including transaction amount, product information, transaction entity, transaction time, transaction country or region, and payment status, generates an order dataset, obtains exchange rates in real time by connecting to an external financial institution interface, and converts the order amount into the settlement currency amount according to the exchange rate to obtain the exchange rate conversion result. Based on the order dataset and exchange rate conversion results, the comprehensive taxes payable are calculated by connecting to an external tax policy interface, and the tax amount is obtained. Based on the order dataset, exchange rate conversion results, and tax amount, a natural language rule conversion algorithm is invoked to parse the input revenue sharing rules into structured executable rules. The natural language rule conversion algorithm uses an improved bidirectional long short-term memory network with dispute memory gating to extract features from the positive and negative context information, obtain the positive and negative hidden state sequences, and combine them with historical transaction data to dynamically optimize the revenue sharing parameters through a deep deterministic strategy gradient algorithm to generate the revenue sharing rules applicable to the current order. Based on the order dataset, currency conversion results, tax amount and revenue sharing rules, the order status and risk characteristics are judged, and if the revenue sharing conditions are met, a revenue sharing calculation instruction is generated. In response to the revenue sharing calculation instruction, calculate the revenue sharing amount for each entity based on the exchange rate conversion result, tax amount, and revenue sharing rules; Verify whether the sum of the accounts of each entity is equal to the settlement currency amount after deducting taxes and fees. If the verification passes, generate a settlement execution instruction. Based on the settlement execution instruction, the amount of the settlement is transferred to the corresponding entity account by connecting to the interface of the external payment institution, the settlement result is obtained, and the arrival time and status are recorded; The order dataset, currency conversion results, tax amount, revenue sharing rules, revenue sharing amount, settlement results, and related logs are written into the blockchain and relational database.

[0008] Furthermore, the step of parsing the input revenue sharing rules into structured executable rules includes: The input revenue sharing rule text is segmented and encoded to obtain the initial feature sequence; The initial feature sequence is input into a bidirectional long short-term memory network with dispute memory gating, and features are extracted from the positive and negative context information respectively to obtain the positive and negative hidden state sequences, and the context state vector at each position is obtained. The context state vector at each location is fused with global rule context embedding information to form an enhanced feature representation; The enhanced feature representation is input into the conditional random field layer, and then jointly decoded through the globally normalized transition probability matrix to output the entity category label and logical relationship label at each position in sequence. Based on entity category tags and logical relationship tags, a structured template mapping is used to generate a revenue sharing rule tree structure; The revenue sharing rule tree structure is standardized and translated to generate structured executable rule expressions.

[0009] Furthermore, the feature extraction operation of the bidirectional long short-term memory network with dispute memory gating specifically includes: During model training and inference, the historical dispute sample library of revenue sharing rules is called in real time. The historical dispute sample library of revenue sharing rules is constructed by collecting and organizing historical orders and corresponding revenue sharing rule texts that have occurred in the system, and performing the same word segmentation and encoding. For each initial feature sequence, matching is performed based on cosine similarity. The historical dispute feature vectors are sorted from high to low according to the similarity score, and the one with the highest similarity score is selected as the most relevant historical dispute feature vector. In the gating mechanism of the bidirectional long short-term memory network unit, a dispute memory gating parameter is added. Specifically, the hidden state of the current input initial feature sequence and the most relevant historical dispute feature vector are input into a learnable gating layer for fusion. The gating layer outputs a dispute memory gating parameter between 0 and 1 based on the similarity score between the current initial feature sequence and the historical dispute features. According to the dispute memory gating parameter, the hidden state of the current initial feature sequence and the most relevant historical dispute feature vector are weighted and fused into the hidden state of the new initial feature sequence. The hidden state of the new initial feature sequence is used as one of the inputs to the activation values ​​of the subsequent forget gate, input gate and output gate. The activation values ​​of each gate are calculated according to the standard long short-term memory network gating structure to generate the fused context state vector. The fused context state vector is aggregated using pooling operations, specifically including max pooling and average pooling of all hidden states in the context state vector. The resulting pooled vectors are then concatenated to form a fixed-length global feature vector, which serves as the enhanced feature representation.

[0010] Furthermore, the step of dynamically optimizing the revenue sharing rule parameters using a deep deterministic strategy gradient algorithm includes: A state vector is constructed using the current revenue sharing rule parameters, historical transaction volume within a fixed period, number of disputes, and average settlement period. The revenue sharing rule parameters are a set of parameters extracted from a structured executable rule expression, specifically including the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. Set the action space. The action is to apply a numerical adjustment to each of the four parameters: platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The adjustment step size and boundary are set when the parameters are initialized. Define a reward function where the reward value equals the percentage of dispute-free orders multiplied by its weight, plus the settlement cycle optimization rate multiplied by its weight. The weights of the percentage of dispute-free orders and the settlement cycle optimization rate are both positive and sum to 1. The percentage of dispute-free orders is equal to the number of dispute-free orders in the current period divided by the total number of orders, and the settlement cycle optimization rate is equal to the negative of the difference between the average settlement cycle in the current period and the industry average settlement cycle. Based on the current state vector and action space, a deep deterministic policy gradient algorithm is adopted to generate parameter adjustment actions using a parameterized policy network, and output the adjustment amount of the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The optimization iteration operation specifically includes: (1) Apply the parameter adjustment action output by the strategy network to the current revenue sharing rule parameters to generate a new set of revenue sharing rule parameters, which is then immediately used for the revenue sharing settlement execution of subsequent actual orders; (2) Based on the actual order settlement results, collect the proportion of dispute-free orders and the settlement cycle optimization rate in the new cycle as reward signals, calculate the corresponding reward value for this cycle according to the defined reward function, and combine the current state vector, the executed parameter adjustment actions, the reward value and the newly collected state vector in the next cycle as a set of training samples and store them in the experience playback buffer. (3) During the network parameter training phase, a batch of random samples of state-action-reward-next state samples are taken from the experience replay buffer. The optimization objective is to maximize the cumulative long-term reward value. (4) After each round of optimization iteration, the new set of revenue sharing rules parameters is continuously used for the settlement and result feedback of the next batch of actual orders. Steps (1) to (3) are continuously repeated until the evaluation index meets the termination condition. Then, the current optimal set of revenue sharing rules parameters is determined as the revenue sharing rule applicable to the current order.

[0011] Furthermore, the deep deterministic strategy gradient algorithm includes an exploration strategy guided by prior knowledge of revenue sharing anomalies, specifically including: Before system deployment, all the accounting parameter configurations in the historical orders of the statistics platform that have caused accounting disputes or fund recovery are used to form a set of accounting anomaly prior parameters; During algorithm training, the exploration strategy is initialized based on the set of prior parameters for revenue sharing anomalies, so that the strategy network prioritizes sampling the range of anomaly parameters and the policy noise sampling weights are adjusted. After each parameter adjustment action is generated, it is determined whether it falls within the interval covered by the set of prior parameters for revenue sharing anomalies. If it does, the sampling probability of the interval is reduced and a penalty term is added to the interval action. Based on newly emerging revenue sharing dispute data, the set of prior parameters for revenue sharing anomalies is dynamically supplemented.

[0012] Furthermore, the steps for generating the revenue sharing calculation instruction specifically include: Verify the current order status to determine if the goods have been received and there is no refund record; Collect the product category of the order, the buyer's historical refund rate, the logistics delivery time and the order amount to form an order feature vector; The order feature vector is input into a pre-trained gradient boosting tree model, and then processed through multiple decision trees to obtain the output of each decision tree. Based on the output and weight of each decision tree, the weighted sum is added and the model bias term is applied. After passing through the Sigmoid activation function, the order refund probability is obtained. If the probability of order refund is less than or equal to 5%, a revenue sharing calculation instruction will be generated immediately, and the revenue sharing execution process will begin. If the probability of order refund is greater than 15%, the triggering of the revenue sharing calculation instruction will be postponed until the order exceeds the platform's stipulated refund time limit, and risk warning information will be pushed to the user's terminal.

[0013] Furthermore, the revenue sharing amounts for each entity include the platform's revenue sharing amount, the supplier's revenue sharing amount, and the distributor's revenue sharing amount.

[0014] Furthermore, the external financial institution interface connects to at least two external financial institutions' RESTful APIs, retrieves real-time exchange rate data once per second, and caches it in a local database.

[0015] The beneficial effects of this invention are: This invention significantly improves the efficiency and accuracy of revenue sharing and settlement on cross-border B2B2C platforms by introducing smart contracts and automated calculation modules. The system can preset and flexibly configure multi-level revenue sharing rules, automatically completing complex revenue sharing logic and amount allocation, increasing revenue sharing efficiency by 90% compared to manual methods and reducing the error rate to below 0.1%. Simultaneously, it integrates currency conversion and multi-country tax calculation functions, connects to external data sources in real time, and automates the entire process of currency exchange and tax accounting, shortening the settlement cycle from the traditional 7-15 days to 1-2 days, significantly reducing supplier capital occupation and business risks.

[0016] This invention utilizes blockchain technology to immutably record the entire transaction and revenue sharing process, ensuring real-time traceability of revenue sharing rules, settlement details, and fund flows, thus greatly enhancing the system's security and compliance. The system also supports customized complex revenue sharing schemes (such as tiered commission rates), multi-currency settlement, and adaptation to multiple national tax systems, enabling flexible responses to various cross-border business scenarios and diverse transaction needs, significantly enhancing the intelligence and standardization of cross-border supply chain finance and fund management. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0018] Figure 1 This is a schematic diagram of a cross-border B2B2C revenue sharing and settlement system based on reinforcement learning proposed in this invention; Figure 2 This is a flowchart of a cross-border B2B2C revenue sharing and settlement system based on reinforcement learning proposed in this invention; Figure 3 This is a data flow diagram of a cross-border B2B2C revenue sharing and settlement system based on reinforcement learning proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1-3 A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning includes the following modules: The transaction information collection module is used to acquire cross-border order information and generate an order dataset. The order information includes transaction amount, product information, transaction entity, transaction time, transaction country or region, and payment status. The currency conversion module is used to obtain exchange rate data in real time by connecting to the interfaces of at least two external financial institutions, and convert the order amount into the settlement currency amount according to the exchange rate to obtain the currency conversion result; The tax calculation module is used to connect to external tax policy interfaces, calculate the comprehensive taxes payable based on the order dataset and exchange rate conversion results, and obtain the tax amount. The smart contract module is used to call the natural language rule conversion algorithm to parse the input revenue sharing rule text into a structured executable rule expression, and combine it with historical transaction data to dynamically optimize the revenue sharing rule parameters through a deep deterministic strategy gradient algorithm to generate the revenue sharing rule parameters applicable to the current order. The revenue sharing rule parameters include the platform revenue sharing ratio, the supplier revenue sharing ratio, the distributor commission ratio, and the distributor tiered commission threshold. The order risk assessment module is used to judge the order status and risk characteristics based on the order dataset, exchange rate conversion results, tax amount and revenue sharing rules. It uses a gradient boosting tree model to calculate the order refund probability based on features such as order product category, buyer's historical refund rate, logistics delivery time and order amount, and determines the revenue sharing trigger time according to a preset threshold, and generates revenue sharing calculation instructions. The revenue sharing calculation module is used to respond to revenue sharing calculation instructions. Based on the exchange rate conversion results, tax amount and revenue sharing rules, it automatically calculates the platform revenue sharing amount, supplier revenue sharing amount and distributor revenue sharing amount, and verifies the sum of the revenue sharing amounts of each entity. If the verification passes, it generates a settlement execution instruction. The settlement module is used to transfer the split amount to the corresponding entity account based on the settlement execution instruction and by connecting to the interface of the external payment institution, and to record the arrival time and settlement status; The data storage and traceability module adopts a hybrid architecture of blockchain and relational database to store order datasets, currency conversion results, tax amounts, revenue sharing rules, revenue sharing amounts, settlement results, and related logs, achieving full-process data immutability and multi-dimensional traceability.

[0021] In this embodiment, the modules are interconnected using the following method: Obtain cross-border order information, generate order datasets, obtain exchange rates in real time by connecting to the RESTful APIs of at least two external financial institutions, and convert the order amount into the settlement currency amount according to the exchange rate updated every second, using high-precision floating-point arithmetic to obtain the exchange rate conversion results; Based on the order dataset and currency conversion results, by connecting to an external tax policy interface, and combining the commodity code, transaction country or region, and the identities of the buyer and seller, the tax rate is automatically adapted to calculate the comprehensive tax payable and obtain the tax amount. Based on the order dataset, currency conversion results, and tax amount, a natural language rule conversion algorithm using word segmentation, encoding, bidirectional long short-term memory network, and conditional random field structure is invoked to parse the input revenue sharing rules into structured executable rules. Historical transaction data includes transaction volume, number of disputes, and average settlement cycle in the past 30 days. Through a deep deterministic strategy gradient algorithm, the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and tiered commission threshold are dynamically optimized with the goal of optimizing the proportion of dispute-free orders and the settlement cycle optimization rate, thereby generating the revenue sharing rules applicable to the current order. Based on the order dataset, exchange rate conversion results, tax amount and revenue sharing rules, the order status and risk characteristics are judged. The gradient boosting tree model is used to predict the order refund probability based on features such as product category, buyer's historical refund rate, logistics receipt time and order amount. If the order status is "received without refund" and the order refund probability is less than or equal to 5%, a revenue sharing calculation instruction is generated. In response to the revenue sharing calculation instruction, the revenue sharing amount of each entity is automatically calculated based on the exchange rate conversion result, tax amount and structured executable revenue sharing rules, supporting complex rules such as tiered commission, condition triggering, and floating ratio; The system verifies whether the sum of the accounts of each entity is equal to the settlement currency amount after deducting taxes and fees. If the verification fails, the system automatically terminates the current settlement process and outputs an alarm log. If the verification passes, a settlement execution instruction is generated. Based on the settlement execution instruction, the amount of the settlement is transferred to the corresponding entity account by connecting to the interface of the external payment institution. It supports parallel transfers through multiple channels and records the arrival time and status. The order dataset, currency conversion results, tax amounts, revenue sharing rules, revenue sharing amounts, settlement results, and related logs are written into a blockchain and a relational database. The blockchain ensures the data is immutable, and the relational database supports multi-dimensional retrieval and auditing across the entire order, settlement, and revenue sharing chain. In this embodiment, the cross-border order information specifically includes transaction amount, product information, transaction entity, transaction time, transaction country or region, and payment status.

[0022] In this embodiment, the step of parsing the input revenue sharing rules into structured executable rules includes: The input revenue sharing rule text is segmented and encoded to obtain an initial feature sequence. The segmentation method is based on a domain dictionary, and the encoding method combines word vectors and business tag embedding. The initial feature sequence is input into a bidirectional long short-term memory network with dispute memory gating. Feature extraction is performed on the positive and negative context information respectively to obtain the hidden state sequences of the positive and negative sides, and the context state vector at each position is obtained. The dispute memory gating mechanism dynamically adjusts the gating parameters based on the similarity with the features of historical dispute rules. The context state vector at each position is fused with global rule context embedding information to form an enhanced feature representation, wherein the global context embedding is generated by combining the order context and the rule history sample pool; The enhanced feature representation is input into the conditional random field layer, and then jointly decoded through the globally normalized transition probability matrix. The entity category label and logical relationship label at each position are output sequentially. The label system includes business elements such as revenue sharing participants, revenue sharing ratio, tiered threshold, and triggering conditions. Based on entity category tags and logical relationship tags, a structured template mapping is used to generate a revenue sharing rule tree structure. The template mapping rules are automatically loaded according to industry business specifications and platform configuration. The revenue sharing rule tree structure is normalized and translated to generate a structured executable rule expression, which serves as the standard input for subsequent revenue sharing amount calculation and automatic contract execution.

[0023] In this embodiment, the feature extraction operation of the bidirectional long short-term memory network with dispute memory gating specifically includes: During model training and inference, the historical dispute sample library of revenue sharing rules is called in real time. The historical dispute sample library of revenue sharing rules is constructed by collecting and organizing historical orders and corresponding revenue sharing rule texts that have occurred in the system, and performing the same word segmentation and encoding. The word segmentation and encoding method is consistent with the main process, and each dispute sample is labeled with dispute type and processing result. For each initial feature sequence, matching is performed based on cosine similarity. The historical dispute feature vectors are sorted from high to low according to the similarity score, and the one with the highest similarity score is selected as the most relevant historical dispute feature vector. The similarity is measured by the cosine angle between the initial feature sequence and the historical dispute feature vector. In the gating mechanism of the bidirectional long short-term memory network unit, a dispute memory gating parameter is added. Specifically, the hidden state of the current input initial feature sequence and the most relevant historical dispute feature vector are input into a learnable gating layer for fusion. The gating layer outputs a dispute memory gating parameter between 0 and 1 based on the similarity score between the current initial feature sequence and the historical dispute features. During fusion, a weighted linear combination method is adopted, that is, the dispute memory gating parameter is used as the weight, and the current hidden state and the historical dispute feature vector are multiplied by the weight and its complement (1 minus the weight) respectively. The weighted results are then added together to form the hidden state of the new initial feature sequence. The hidden state of the new initial feature sequence is used as one of the inputs to the activation values ​​of the subsequent forget gate, input gate and output gate. The activation values ​​of each gate are calculated according to the standard long short-term memory network gating structure to generate the fused context state vector. The calculation of the activation values ​​of each gate adopts the joint linear transformation of the current hidden state, historical dispute features and input features plus the sigmoid activation function. The fused context state vector is aggregated using pooling operations, specifically including maximum pooling and average pooling of all hidden states in the context state vector. The resulting pooled vectors are then concatenated to form a fixed-length global feature vector, which serves as an enhanced feature representation. The pooling dimension and concatenation order are configured according to business requirements.

[0024] In this embodiment, the step of dynamically optimizing the revenue sharing rule parameters using a deep deterministic strategy gradient algorithm includes: A state vector is constructed using the current revenue sharing rule parameters, transaction volume within a historical fixed period, number of disputes, and average settlement period. The revenue sharing rule parameters are a set of parameters extracted from a structured executable rule expression, specifically including the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The historical fixed period is generally set to 30 days. Set the action space. The action is to apply a numerical adjustment to each of the four parameters: platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The adjustment step size and boundary are set when the parameters are initialized. All parameter adjustments are subject to business constraints. The sum of the revenue sharing ratios does not exceed 1, and the threshold change does not exceed the preset extreme value. Define a reward function where the reward value equals the percentage of dispute-free orders multiplied by its weight, plus the settlement cycle optimization rate multiplied by its weight. The weights for both the percentage of dispute-free orders and the settlement cycle optimization rate are positive and sum to 1. The percentage of dispute-free orders is equal to the number of dispute-free orders in the current period divided by the total number of orders. The settlement cycle optimization rate is equal to the negative of the difference between the average settlement cycle in the current period and the industry average settlement cycle. The weight parameters and industry averages in the reward function can be dynamically adjusted according to the actual business scenario. Based on the current state vector and action space, a deep deterministic policy gradient algorithm is adopted. The parameterized policy network is used to generate parameter adjustment actions and output the adjustment amount of the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio and distributor tiered commission threshold. The policy network and value network are both multilayer perceptron structures, and the input and output dimensions correspond one-to-one with the revenue sharing business parameters. The optimization iteration operation specifically includes: (1) Apply the parameter adjustment action output by the strategy network to the current revenue sharing rule parameters to generate a new revenue sharing rule parameter set, which is immediately used for the revenue sharing settlement execution of subsequent actual orders. All parameter adjustment processes and results are automatically logged. (2) Based on the actual order settlement results, collect the proportion of dispute-free orders and the settlement cycle optimization rate in the new cycle as reward signals, calculate the corresponding reward value for this cycle according to the defined reward function, and combine the current state vector, the executed parameter adjustment actions, the reward value and the newly collected state vector in the next cycle as a set of training samples and store them in the experience playback buffer. The capacity of the experience playback buffer and the sampling strategy can be dynamically expanded according to the business pressure. (3) During the network parameter training phase, a batch of random samples of state-action-reward-next state samples are taken from the experience replay buffer to update the parameters of the policy network and the value network through gradient backpropagation. The optimization objective is to maximize the cumulative long-term reward value. During the optimization process, the legality constraint of the revenue sharing parameter is introduced to prevent abnormal output. (4) After each round of optimization iteration, the new set of revenue sharing rules parameters is continuously used for the settlement and result feedback of the next batch of actual orders. Steps (1) to (3) are continuously repeated until the evaluation index meets the termination condition. Then, the current optimal set of revenue sharing rules parameters is determined as the revenue sharing rule applicable to the current order.

[0025] In this embodiment, the deep deterministic strategy gradient algorithm includes an exploration strategy guided by prior knowledge of revenue sharing anomalies, specifically including: Before system deployment, all the revenue sharing parameter configurations that have caused revenue sharing disputes or fund recovery in the historical orders of the statistical platform are counted to form a set of revenue sharing abnormal prior parameters. The revenue sharing parameter configurations include the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio and tiered commission threshold, and the abnormal sample intervals are grouped and marked according to the actual dispute type. During the algorithm training process, the exploration strategy is initialized based on the set of prior parameters for revenue sharing anomalies. This allows the strategy network to prioritize sampling the range of anomalies and adjust the sampling weight of strategy noise. During the sampling process, business risk weight factors are added to actions that fall into the anomaly range. After each parameter adjustment action is generated, it is determined whether it falls within the interval covered by the set of prior parameters for revenue sharing anomalies. If it does, the sampling probability of the interval is reduced and a penalty term is added to the interval action. The penalty term can be a negative weighted sum of the reward function or a direct increase in the regularization loss of the corresponding parameter adjustment. Based on newly emerging revenue sharing dispute data, the system dynamically supplements the set of prior parameters for revenue sharing anomalies. The system periodically and automatically updates the prior set and synchronizes it to the policy network exploration distribution, so that the anomaly prior policy can continuously adapt to changes in the business environment.

[0026] In this embodiment, the step of generating the revenue sharing calculation instruction specifically includes: Verify the current order status to determine whether the goods have been received and there is no refund record. The order status data is automatically synchronized by the order management system and the data integrity is verified. The product category, buyer's historical refund rate, logistics delivery time, and order amount of the order are collected to form an order feature vector. The feature data is taken from the order and user behavior database and standardized into a unified input format. The order feature vector is input into a pre-trained gradient boosting tree model, and the model is calculated sequentially through multiple decision trees to obtain the output of each decision tree. The model parameters and structure are obtained through offline training based on the platform's historical order refund data. The number of decision trees and the maximum depth are set according to business needs. Based on the output and weight of each decision tree, the weighted sum is added and the model bias term is applied. After passing through the Sigmoid activation function, the order refund probability is obtained. The weights and bias term are dynamically adjusted as the model is trained. The Sigmoid function outputs the probability in the range of 0 to 1. If the probability of order refund is less than or equal to 5%, a revenue sharing calculation instruction will be generated immediately, and the revenue sharing execution process will begin. If the probability of order refund is greater than 15%, the triggering of the revenue sharing calculation instruction will be postponed until the order exceeds the platform's stipulated refund time limit, and risk warning information will be pushed to the user's terminal. The platform supports manual review or dynamic reassessment of such high-risk orders.

[0027] In this embodiment, the revenue sharing amount for each entity includes the platform revenue sharing amount, the supplier revenue sharing amount, and the distributor revenue sharing amount. All revenue sharing amounts are automatically calculated by structured executable rule expressions, accurate to the cent level, and are settled according to the respective entity accounts.

[0028] In this embodiment, the external financial institution interface connects to at least two external financial institutions' RESTful APIs, retrieves real-time exchange rate data once per second, and caches it in a local database. The exchange rate data includes currency pairs, the latest exchange rate value, and data timestamps. The interface supports failover and fault tolerance, and all historical exchange rate snapshots can be used for post-event auditing and backtracking.

[0029] Example 1: To verify the feasibility of this invention in practice, it was applied to a cross-border B2B2C e-commerce platform, and the application process and effects of the technical solution of this invention in actual business were explained in detail.

[0030] A platform received a cross-border order in June 2025 involving buyer A (China), supplier B (Germany), and distributor C (Singapore). Order details are as follows: Order Number: ORD202506001; Product Category: Home Appliances; Order Amount: USD 1000; Order Currency: USD; Order Placement Time: 2025-06-01 10:00:00; Payment Status: Paid; Logistics Status: Received; Buyer Historical Refund Rate: 3%; Logistics Delivery Time: 7 days.

[0031] The system obtains real-time exchange rate data for USD to RMB, EUR, and SGD once per second through the interfaces of the connected external financial institutions (Bank A and Bank B) and caches it in the local database.

[0032] The latest exchange rates for the day (the system automatically verifies the consistency between the two exchange rates) are: USD / CNY = 7.1973; USD / EUR = 0.8806; USD / SGD = 1.2894. The settlement currency for this order is Chinese Yuan (CNY). The system automatically converts the order amount according to the real-time exchange rate: Order settlement amount = 1000 × 7.1973 = 7,197.30 yuan.

[0033] The system automatically reads the tax rates for household appliances in Sino-German cross-border transactions by connecting to an external tax policy interface. Commodity code (HS code): 850940; Customs duty: 10%; Value-added tax (VAT): 13%. Calculations show: Customs duty = 7,197.30 × 10% = 719.73 yuan; VAT = 7,197.30 × 13% = 935.65 yuan; Total tax = 719.73 + 935.65 = 1,655.38 yuan. Administrator input rules: "Supplier 80%, Distributor 10%, Platform 10%. If a distributor's monthly sales exceed 50,000 yuan, the commission will increase by 5%."

[0034] The system uses word segmentation, word vector encoding, bidirectional long short-term memory network (Bi-LSTM) and conditional random field (CRF) to parse the rule into a structured expression, and combines the transaction volume of similar orders in the past 30 days, the average settlement cycle (5 days), and the number of disputes (2 times) to automatically optimize the parameters using the deep deterministic policy gradient (DDPG) algorithm.

[0035] Table 1 Order Revenue Sharing Parameters

[0036] As shown in Table 1, distributors' commissions increased due to exceeding monthly sales targets, while the platform's commission rate decreased slightly.

[0037] The system collects information such as product category, buyer's historical refund rate, logistics delivery time, and order amount. This information is then input into the trained GBDT model. The model outputs that the refund probability for this order is 3.6% (less than 5%), and the order status is "received with no refund". The system automatically triggers the revenue sharing calculation instruction.

[0038] The base amount for revenue sharing is 7,197.30 - 1,655.38 = 5,541.92 yuan. Based on the optimized ratio (rounded to the nearest cent): Platform: 5,541.92 × 9% = 499.76 yuan; Supplier: 5,541.92 × 80% = 4,433.54 yuan; Distributor: 5,541.92 × 11% = 609.61 yuan. Due to potential minor discrepancies in the total after rounding, the system automatically allocates any fractional yuan to the platform account.

[0039] The system collects order product category, buyer's historical refund rate (3%), logistics delivery time (7 days), and order amount (7,197.30 yuan), inputs these data into a trained GBDT model, and outputs a refund probability of 3.6% (less than 5%). The order status is "Received, No Refund," and the system automatically triggers a revenue sharing calculation instruction. The revenue sharing amount is automatically calculated according to the optimized ratio: Platform revenue sharing amount = 499.76 yuan; Supplier revenue sharing amount = 4,433.54 yuan; Distributor revenue sharing amount = 609.61 yuan; Total = 499.76 + 4,433.54 + 609.61 = 5,542.91 yuan; Verification: Settlement base amount = 7,197.30 - 1,655.38 = 5,541.92 yuan. Verification error = 5,542.91 - 5,541.92 = 0.99 yuan (the system automatically allocates the error to the platform).

[0040] The system calls the payment institution's interface to automatically transfer the various revenue amounts to the platform, supplier, and distributor accounts. The arrival time and status are automatically written to the blockchain and database. The entire process retains operation logs and revenue details, making it auditable and traceable.

[0041] Table 2 Order Revenue Sharing Optimization and Settlement Details

[0042] As shown in Table 2, after AI optimization in this embodiment, the platform's revenue sharing ratio decreased from 10% to 9%, the distributor's ratio increased from 10% to 11% due to improved sales performance, and the supplier's ratio remained unchanged at 80%. The revenue sharing amounts for all parties have been automatically settled. The error sharing mechanism in the table ensures that the total revenue sharing amount is strictly aligned with the settlement base amount, realizing intelligent adjustment of the revenue sharing process, business incentives, and a balance of interests among multiple parties.

[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning, characterized in that, Includes the following modules: The transaction information collection module is used to acquire cross-border order information and generate order datasets; The currency conversion module is used to obtain exchange rate data in real time through the interface of external financial institutions, and convert the order amount into the settlement currency amount according to the exchange rate to obtain the currency conversion result; The tax calculation module is used to connect to external tax policy interfaces, calculate the comprehensive taxes payable based on the order dataset and exchange rate conversion results, and obtain the tax amount. The smart contract module is used to call the natural language rule conversion algorithm to parse the input revenue sharing rule text into a structured executable rule expression. The natural language rule conversion algorithm adopts an improved bidirectional long short-term memory network with dispute memory gating, extracts features from the positive and negative context information respectively, obtains the positive and negative hidden state sequences, and combines historical transaction data to dynamically optimize the revenue sharing rule parameters through a deep deterministic strategy gradient algorithm to generate the revenue sharing rule parameters applicable to the current order. The order risk assessment module is used to determine the order status and risk characteristics. It adopts a gradient boosting tree model to determine the order status and risk characteristics. If the revenue sharing conditions are met, a revenue sharing calculation instruction is generated. The revenue sharing calculation module is used to respond to revenue sharing calculation instructions, calculate the revenue sharing amount of each entity based on the exchange rate conversion result, tax amount and revenue sharing rules, and verify the sum of the revenue sharing amounts of each entity. If the verification passes, a settlement execution instruction is generated. The settlement module is used to connect to the interface of external payment institutions, transfer the split amount to the corresponding entity account, and record the arrival time and settlement status; The data storage and traceability module is used to store order datasets, currency conversion results, tax amounts, revenue sharing rules, revenue sharing amounts, settlement results, and related logs.

2. The cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 1, characterized in that, The modules are connected in the following way: The system obtains cross-border order information, including transaction amount, product information, transaction entity, transaction time, transaction country or region, and payment status, generates an order dataset, obtains exchange rates in real time by connecting to an external financial institution interface, and converts the order amount into the settlement currency amount according to the exchange rate to obtain the exchange rate conversion result. Based on the order dataset and exchange rate conversion results, the comprehensive taxes payable are calculated by connecting to an external tax policy interface, and the tax amount is obtained. Based on the order dataset, exchange rate conversion results, and tax amount, a natural language rule conversion algorithm is invoked to parse the input revenue sharing rules into structured executable rules. The natural language rule conversion algorithm uses an improved bidirectional long short-term memory network with dispute memory gating to extract features from the positive and negative context information, obtain the positive and negative hidden state sequences, and combine them with historical transaction data to dynamically optimize the revenue sharing parameters through a deep deterministic strategy gradient algorithm to generate the revenue sharing rules applicable to the current order. Based on the order dataset, exchange rate conversion results, tax amount and revenue sharing rules, the order status and risk characteristics are judged, and if the revenue sharing conditions are met, a revenue sharing calculation instruction is generated. In response to the revenue sharing calculation instruction, calculate the revenue sharing amount for each entity based on the exchange rate conversion result, tax amount, and revenue sharing rules; Verify whether the sum of the accounts of each entity is equal to the settlement currency amount after deducting taxes and fees. If the verification passes, generate a settlement execution instruction. Based on the settlement execution instruction, the amount of the settlement is transferred to the corresponding entity account by connecting to the interface of the external payment institution, the settlement result is obtained, and the arrival time and status are recorded; The order dataset, currency conversion results, tax amount, revenue sharing rules, revenue sharing amount, settlement results, and related logs are written into the blockchain and relational database.

3. The cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, The step of parsing the input revenue sharing rules into structured executable rules includes: The input revenue sharing rule text is segmented and encoded to obtain the initial feature sequence; The initial feature sequence is input into a bidirectional long short-term memory network with dispute memory gating, and features are extracted from the positive and negative context information respectively to obtain the positive and negative hidden state sequences, and the context state vector at each position is obtained. The context state vector at each location is fused with global rule context embedding information to form an enhanced feature representation; The enhanced feature representation is input into the conditional random field layer, and then jointly decoded through the globally normalized transition probability matrix to output the entity category label and logical relationship label at each position in sequence. Based on entity category tags and logical relationship tags, a structured template mapping is used to generate a revenue sharing rule tree structure; The revenue sharing rule tree structure is standardized and translated to generate structured executable rule expressions.

4. The cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 3, characterized in that, The feature extraction operation of the bidirectional long short-term memory network with dispute memory gating specifically includes: During model training and inference, the historical dispute sample library of revenue sharing rules is called in real time. The historical dispute sample library of revenue sharing rules is constructed by collecting and organizing historical orders and corresponding revenue sharing rule texts that have occurred in the system, and performing the same word segmentation and encoding. For each initial feature sequence, matching is performed based on cosine similarity. The historical dispute feature vectors are sorted from high to low according to the similarity score, and the one with the highest similarity score is selected as the most relevant historical dispute feature vector. In the gating mechanism of the bidirectional long short-term memory network unit, a dispute memory gating parameter is added. Specifically, the hidden state of the current input initial feature sequence and the most relevant historical dispute feature vector are input into a learnable gating layer for fusion. The gating layer outputs a dispute memory gating parameter between 0 and 1 based on the similarity score between the current initial feature sequence and the historical dispute features. According to the dispute memory gating parameter, the hidden state of the current initial feature sequence and the most relevant historical dispute feature vector are weighted and fused into the hidden state of the new initial feature sequence. The hidden state of the new initial feature sequence is used as one of the inputs to the activation values ​​of the subsequent forget gate, input gate and output gate. The activation values ​​of each gate are calculated according to the standard long short-term memory network gating structure to generate the fused context state vector. The fused context state vector is aggregated using pooling operations, specifically including max pooling and average pooling of all hidden states in the context state vector. The resulting pooled vectors are then concatenated to form a fixed-length global feature vector, which serves as the enhanced feature representation.

5. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, The steps for dynamically optimizing the revenue sharing rule parameters using a deep deterministic strategy gradient algorithm include: A state vector is constructed using the current revenue sharing rule parameters, historical transaction volume within a fixed period, number of disputes, and average settlement period. The revenue sharing rule parameters are a set of parameters extracted from a structured executable rule expression, specifically including the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. Set the action space. The action is to apply a numerical adjustment to each of the four parameters: platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The adjustment step size and boundary are set when the parameters are initialized. Define a reward function where the reward value equals the percentage of dispute-free orders multiplied by its weight, plus the settlement cycle optimization rate multiplied by its weight. The weights of the percentage of dispute-free orders and the settlement cycle optimization rate are both positive and sum to 1. The percentage of dispute-free orders is equal to the number of dispute-free orders in the current period divided by the total number of orders, and the settlement cycle optimization rate is equal to the negative of the difference between the average settlement cycle in the current period and the industry average settlement cycle. Based on the current state vector and action space, a deep deterministic policy gradient algorithm is adopted to generate parameter adjustment actions using a parameterized policy network, and output the adjustment amount of the platform revenue sharing ratio, supplier revenue sharing ratio, distributor commission ratio, and distributor tiered commission threshold. The optimization iteration operation specifically includes: (1) Apply the parameter adjustment action output by the strategy network to the current revenue sharing rule parameters to generate a new set of revenue sharing rule parameters, which is then immediately used for the revenue sharing settlement execution of subsequent actual orders; (2) Based on the actual order settlement results, collect the proportion of dispute-free orders and the settlement cycle optimization rate in the new cycle as reward signals, calculate the corresponding reward value for this cycle according to the defined reward function, and combine the current state vector, the executed parameter adjustment actions, the reward value and the newly collected state vector in the next cycle as a set of training samples and store them in the experience playback buffer. (3) During the network parameter training phase, a batch of random samples of state-action-reward-next state samples are taken from the experience replay buffer. The optimization objective is to maximize the cumulative long-term reward value. (4) After each round of optimization iteration, the new set of revenue sharing rules parameters is continuously used for the settlement and result feedback of the next batch of actual orders. Steps (1) to (3) are continuously repeated until the evaluation index meets the termination condition. Then, the current optimal set of revenue sharing rules parameters is determined as the revenue sharing rule applicable to the current order.

6. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, Deep deterministic policy gradient algorithms include exploration strategies guided by prior knowledge of revenue sharing anomalies, specifically including: Before system deployment, all the accounting parameter configurations in the historical orders of the statistics platform that have caused accounting disputes or fund recovery are used to form a set of accounting anomaly prior parameters; During algorithm training, the exploration strategy is initialized based on the set of prior parameters for revenue sharing anomalies, so that the strategy network prioritizes sampling the range of anomaly parameters and the policy noise sampling weights are adjusted. After each parameter adjustment action is generated, it is determined whether it falls within the interval covered by the set of prior parameters for revenue sharing anomalies. If it does, the sampling probability of the interval is reduced and a penalty term is added to the interval action. Based on newly emerging revenue sharing dispute data, the set of prior parameters for revenue sharing anomalies is dynamically supplemented.

7. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, The steps for generating the revenue sharing calculation instruction specifically include: Verify the current order status to determine if the goods have been received and there is no refund record; Collect the product category of the order, the buyer's historical refund rate, the logistics delivery time and the order amount to form an order feature vector; The order feature vector is input into a pre-trained gradient boosting tree model, and then processed through multiple decision trees to obtain the output of each decision tree. Based on the output and weight of each decision tree, the weighted sum is added and the model bias term is applied. After passing through the Sigmoid activation function, the order refund probability is obtained. If the probability of order refund is less than or equal to 5%, a revenue sharing calculation instruction will be generated immediately, and the revenue sharing execution process will begin. If the probability of order refund is greater than 15%, the triggering of the revenue sharing calculation instruction will be postponed until the order exceeds the platform's stipulated refund time limit, and risk warning information will be pushed to the user's terminal.

8. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, The revenue sharing amounts for each entity include the platform's revenue sharing amount, the supplier's revenue sharing amount, and the distributor's revenue sharing amount.

9. A cross-border B2B2C revenue sharing and settlement system based on reinforcement learning according to claim 2, characterized in that, The external financial institution interface is connected to at least two external financial institutions' RESTful APIs, which retrieve real-time exchange rate data once per second and cache it in the local database.

Citation Information

Cited By

  • Fund account intelligent compliance management and control method, device and equipment and storage medium

    CN121414363A