Cross-border e-commerce cost fine management method and device based on big data, equipment and medium
By constructing a cost-event knowledge graph for cross-border e-commerce and allocating costs using implicit driving vectors, the problem of tracing implicit costs in cross-border e-commerce has been solved, enabling refined cost management and decision optimization, and improving operational agility and decision quality.
Patent Information
- Application Number
- CN202511704450.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In cross-border e-commerce cost management, the phenomenon of data silos from multiple sources leads to the omission of hidden costs, the lack of accurate basis for cost allocation, and the difficulty in tracing marketing expenses and overseas warehouse demurrage fees. Decisions lack data support, resulting in large deviations in cost accounting, delayed responses to policy risks, and hindering the improvement of refined operational capabilities.
By acquiring multi-source data from cross-border e-commerce businesses, a cost-event knowledge graph is constructed. A pre-trained graph-structured causal exploration model is used to evaluate the long-term marginal increment of unfulfilled orders, resulting in implicit driving vectors and a list of controllable factors. Cost allocation is then performed based on the knowledge graph and implicit driving vectors, and management decisions are optimized by combining macro information.
It enables refined cost management in cross-border e-commerce operations, improves the transparency and interpretability of cost analysis, enhances the causal basis of business decisions, and improves operational agility and decision-making quality.
Smart Images

Figure CN121563602A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data analysis technology, and in particular relates to a method, device, equipment and medium for refined cost management of cross-border e-commerce based on big data. Background Technology
[0002] As the cross-border e-commerce industry advances towards large-scale and globalization, cost management technologies have emerged, enabling real-time processing and correlation mining of data across the entire process, including procurement, logistics, and compliance, through the integration of multi-source data.
[0003] In traditional methods, cost accounting is mostly based on the order amount and allocated proportionally, with procurement, logistics, marketing and other expenses being split into each order at a fixed ratio; relying on manual processing of paper documents such as customs declarations and logistics bills, and manually calculating the cost per transaction.
[0004] However, the above methods result in data silos from multiple sources, with no correlation between order, warehousing, and customs data. This leads to the omission of hidden costs, a lack of accurate basis for cost allocation, and difficulty in tracing marketing expenses and overseas warehouse demurrage fees to specific orders or SKUs (Stock Keeping Units). It also makes it impossible to quantify the long-term cost impact of factors such as carrier delays and policy adjustments. Decision-making lacks data support, causing companies to frequently face dilemmas such as excessive cost accounting deviations and delayed responses to policy risks, thus hindering the improvement of refined operational capabilities. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, device, equipment, and medium for refined cost management of cross-border e-commerce based on big data, which can accurately allocate costs and analyze hidden costs, in order to address the above-mentioned technical problems.
[0006] Firstly, this application provides a method for refined cost management in cross-border e-commerce based on big data, including:
[0007] Acquire multi-source data from cross-border e-commerce business and preprocess the multi-source data to obtain a business database; multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing invoices, foreign exchange data, and policy data;
[0008] A cost-event knowledge graph is constructed based on the business database. The cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes.
[0009] Based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph to obtain implicit driving vectors and a list of controllable factors.
[0010] Based on the cost-event knowledge graph and implicit driving vectors, the cost of each order is allocated according to the actual event activities, resulting in a cost waterfall for each order;
[0011] The cost waterfall is revised based on macroeconomic information, and management decisions are optimized based on the revised cost waterfall to obtain a business management solution. Macroeconomic information includes, but is not limited to, tariff policies, promotional calendars, and holidays. The business management solution includes executable scheduling of order consolidation, logistics route selection, warehouse allocation, and adjustments to pricing and marketing efforts.
[0012] In one embodiment, multi-source data is preprocessed to obtain a business database, including:
[0013] The business events corresponding to each order are extracted from multi-source data using order identifiers as clues. The extracted business events are then standardized by time zone and timestamp to obtain a time-normalized order event stream.
[0014] Based on preset entity matching rules, the order event stream is parsed and deduplicated using a fuzzy matching algorithm to obtain aligned entity identifiers; entities include, but are not limited to, orders, suppliers, and carriers;
[0015] Based on foreign exchange data from multiple sources, the cost data in the order event stream, which are priced in different currencies, are converted into currency according to the exchange rate on the cost settlement date to obtain a uniformly priced cost amount.
[0016] Based on the product text and attribute information in the SKU data, a pre-trained HS coding annotation model is used to generate multiple corresponding HS coding candidates and confidence scores for each SKU.
[0017] Semantic labels are generated by semantically annotating each order event stream based on policy data and warehousing data.
[0018] Each order event stream is supplemented with fields consisting of entity identifier, uniformly priced fee amount, HS code candidate and its confidence level and semantic label to obtain the business database; the business database includes structured order business data.
[0019] In one embodiment, based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on a cost-event knowledge graph to obtain implicit driving vectors and a list of controllable factors, including:
[0020] Extract business event nodes and business entity nodes related to predefined volatile events from the cost-event knowledge graph;
[0021] By inputting business event nodes and business entity nodes of the same volatile event into a graph-based causal exploration model, the long-term marginal incremental cost and delay probability index of the order corresponding to the volatile event are obtained; the long-term marginal incremental cost includes, but is not limited to, potential return costs and potential warehousing costs; the graph-based causal exploration model is obtained by training causal inference on historical volatile event data;
[0022] Based on the business entity nodes, the long-term marginal incremental cost is decomposed into an interpretable value to obtain the implicit driving vector; the implicit driving vector includes the weight value of each influencing factor.
[0023] Based on the delay probability index, the influencing factors in the implicit driving vector are mapped to items that can be adjusted through operational actions, thus obtaining a list of controllable factors.
[0024] In one embodiment, based on a cost-event knowledge graph and implicit driving vectors, costs are allocated to each order according to actual event activities, resulting in a cost waterfall for each order, including:
[0025] Based on the business event nodes of the cost-event knowledge graph, the entire business process of each order is divided into events to obtain the activity pool corresponding to each order. The activity pool includes procurement, warehousing, storage, picking, packaging, export transportation, customs clearance, overseas warehouse operations, last-mile delivery, return processing, platform settlement and marketing attribution.
[0026] The adjustment value of the driver for the corresponding activity pool is determined based on the weight value of the implicit driving vector; the driver is determined by the cost edge between the business entity node and the business event node.
[0027] Based on the activity pool and drivers, and their corresponding driver and adjustment values, the cost waterfall for each order is obtained;
[0028] The cost waterfall can be obtained using the following formula:
[0029]
[0030] in, The total cost of the cost waterfall; For the activity pool collection, For the first Total cost of multiple orders across multiple activity pools; For orders In the drive The driving value below; For all orders in the first Total drive value under the class activity pool driver; For the first Each activity pool for orders The adjustment value.
[0031] In one embodiment, the method further includes:
[0032] In response to the detection and identification of an order return, refund, or tax refund event, the relevant business event nodes and cost edges are determined by tracing along the association edges of the business entity nodes corresponding to the order based on the cost-event knowledge graph.
[0033] Based on the traced business event nodes and cost edges, the costs previously allocated to the activity pool of the order are reversed according to the preset reversal rules, and the remaining costs after reversal are redistributed to related orders to obtain the updated cost waterfall for related orders.
[0034] In one embodiment, the cost waterfall is revised based on macro-level information, and management decisions are optimized based on the revised cost waterfall to obtain a business management solution, including:
[0035] The cross-border commodity tariff policy text in the policy data is structured and parsed to obtain a set of executable rules; the set of executable rules includes the target country, commodity HS code, value range, and tax calculation rules;
[0036] Based on historical customs clearance data and commodity HS codes, the confidence level of HS code candidates is verified to determine the HS code corresponding to each order.
[0037] Based on the set of executable rules, taxes and fees are calculated according to the HS code and order value of each order.
[0038] Based on promotional calendars and holidays, and according to historical order sequences and historical carrier rates, a pre-trained combined time series forecasting model is used to make predictions, and the order demand, logistics freight rates and return rates within the preset forecast period are obtained.
[0039] The cost waterfall is updated based on taxes, logistics costs, and return rates to obtain the updated cost waterfall.
[0040] Based on the revised cost waterfall and combined with order requirements, management decisions are optimized to obtain a business management solution.
[0041] In one embodiment, a business management solution is obtained by optimizing management decisions based on the revised cost waterfall and order demand, including:
[0042] Construct an optimization objective function that minimizes costs while satisfying time constraints and service level requirements based on the cost waterfall and order demand.
[0043] Based on the objective function, a greedy algorithm or mixed integer programming algorithm is used to solve the optimal solutions for order merging strategy, logistics route selection, warehouse allocation, carrier selection, pricing and marketing adjustment according to the list of controllable factors, so as to obtain the business management solution.
[0044] Secondly, this application also provides a big data-based device for refined cost management in cross-border e-commerce, comprising:
[0045] The business data module is used to acquire multi-source data from cross-border e-commerce business and preprocess the multi-source data to obtain the business database; the multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing invoices, foreign exchange data and policy data;
[0046] The knowledge graph module is used to build a cost-event knowledge graph based on the business database. The cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes.
[0047] The implicit cost-driven discovery module is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph based on historical business data and a pre-trained graph-based causal exploration model to obtain implicit driving vectors and a list of controllable factors.
[0048] The cost analysis module is used to allocate costs for each order based on actual event activities, using a cost-event knowledge graph and implicit driving vectors, to obtain a cost waterfall for each order.
[0049] The management optimization module is used to correct the cost waterfall based on macro information and optimize management decisions based on the corrected cost waterfall to obtain a business management solution. Macro information includes, but is not limited to, tariff policies, promotional calendars, and holidays. The business management solution includes executable scheduling of order merging, logistics route selection, warehouse allocation, and adjustments to pricing and marketing.
[0050] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above-mentioned methods for refined cost management of cross-border e-commerce based on big data.
[0051] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned methods for refined cost management of cross-border e-commerce based on big data.
[0052] The aforementioned methods, devices, equipment, and media for refined cost management in cross-border e-commerce based on big data acquire and preprocess multi-source data, including orders, SKUs, procurement, warehousing, transportation, customs, marketing invoices, foreign exchange, and policies, forming a unified business database. This centralized management of dispersed data resources in cross-border e-commerce operations provides reliable foundational data support for cost calculation and analysis. By constructing a cost-event knowledge graph containing business entity nodes, business event nodes, relational edges, and expense edges, the cost generation mechanism in the cross-border e-commerce business process is visualized and structured, clearly depicting the composition and generation path of various expenses from order creation to fulfillment. Running a pre-trained graph-based causal exploration model on the cost-event knowledge graph quantifies the long-term marginal increment of unfulfilled orders based on historical business data, obtaining implicit driving vectors and a list of controllable factors. This identifies which business factors have a key impact on cost changes, providing causal basis for business decisions. Based on a cost-event knowledge graph and implicit driving vectors, costs are allocated to the actual activities of each order, resulting in an order-level cost waterfall. This allows for a hierarchical and traceable presentation of the cumulative costs of each cross-border e-commerce order throughout the fulfillment process, improving the transparency and interpretability of cost analysis. By using macroeconomic information such as tariff policies, promotional calendars, and holidays as input, the cost waterfall is revised, enabling cost forecasting and allocation to reflect changes in the external environment in real time. Based on the revised results, business management solutions are output, including order consolidation, route selection, warehouse allocation, pricing, and marketing adjustments, thereby improving the agility and decision-making quality of cross-border e-commerce operations. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating the method for refined cost management in cross-border e-commerce based on big data, as described in this invention.
[0055] Figure 2 This is a flowchart illustrating the steps of step S103.
[0056] Figure 3 This is a flowchart illustrating the steps of step S104.
[0057] Figure 4 This is a structural diagram of the cross-border e-commerce cost refinement management device based on big data according to the present invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0059] In one embodiment, such as Figure 1 As shown, a method for refined cost management in cross-border e-commerce based on big data is provided. This embodiment illustrates the method by applying it to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0060] S101. Obtain multi-source data for cross-border e-commerce business and preprocess the multi-source data to obtain a business database; multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing bills, foreign exchange data, and policy data.
[0061] In illustrative terms, multi-source data refers to various heterogeneous data generated during the operation of cross-border e-commerce businesses. This includes: order data (order number, SKU, quantity, transaction price, channel, buyer's country, creation time, etc.); SKU data (product title, attributes, declared value, volume, weight, candidate HS codes, etc.); procurement data (supplier, purchase price, delivery terms, invoice date, etc.); warehousing data (inbound / outbound time, inventory days, occupied volume, etc.); transportation data (waybill number, carrier, route, service level, actual freight, waybill event sequence, etc.); customs data (customs declaration number, declared value, customs duties and value-added tax, HS codes used for customs declaration, etc.); marketing bills (advertising expenses, platform commission, etc.); foreign exchange data (exchange rate time series, settlement currency); and policy data (customs regulations, tariff schedules, processing fee rules, temporary notices, promotional calendars, and statutory holidays, etc.). Multi-source data may include structured tables as well as semi-structured or textual data, such as policy texts and product descriptions.
[0062] Optionally, an ETL / ELT process can be used to extract data from various business systems, such as the platform order database, WMS (Warehouse Management System), TMS (Transportation Management System), financial invoice database, customs declaration system, advertising platform, exchange rate service, and regulatory crawler. This multi-source data is then connected to its corresponding storage components. For example, order / marketing data is written to the data lake in real-time via Kafka, while procurement / policy data is synchronized to HDFS via Spark batch processing. Furthermore, metadata associations are established through the Data Catalog, such as associating orders with carriers and customs declaration numbers, and a unified primary key mapping rule is implemented. For instance, the supplier identifier SP003 is uniformly labeled as a specific supplier. Through this structured integration, a business database is formed.
[0063] S102. Construct a cost-event knowledge graph based on the business database; the cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes.
[0064] A cost-event knowledge graph refers to a graph structure that uses business entities such as orders, goods, suppliers, carriers, and warehouses, as well as business events such as inbound, outbound, shipment, customs clearance, payment collection, and refunds, as graph nodes. Edges represent the relationships between entities or the occurrence relationships between entities and events, and cost information is embedded in the relevant edges as edge weights. It can directly express multi-hop causal links across entities, such as the delays and additional warehousing costs caused by sequentially connecting a supplier, a carrier, a route, and a warehouse. It also facilitates graph structure learning methods in capturing structured spatial correlations. Illustratively, starting from a business database, the graph schema is first defined, including node types, event types, and edge types. For example, nodes are divided into business entity nodes and business event nodes, and edges are divided into relationship edges between entities (e.g., an order indicates that it contains a certain product), and event edges between entities (e.g., an order involves transportation and incurs costs). Costs are stored as edge weights or edge attributes; optionally, costs can be treated as separate nodes. The cost-event knowledge graph visually represents the procurement chain from order to SKU to supplier, the logistics chain from order to overseas warehouse to carrier, and the allocation chain from event to cost to order.
[0065] S103. Based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph to obtain implicit driving vectors and a list of controllable factors.
[0066] Historical business data refers to events that occurred in the past, such as increased return rates and warehouse storage costs due to carrier delays, or increased transportation costs due to supplier delays. A graph-based causal exploration model refers to a model trained offline or pre-trained on the aforementioned historical data and corresponding knowledge graphs. This model can utilize the graph's topology, node / edge attributes, and time information to predict the long-term marginal incremental costs and risks of orders in transit or not yet completed. For example, using the historical business database, if a carrier delay due to a strike occurred in a certain year, difference-difference (DiD) analysis can be used to compare the cost differences before and after the strike. Specifically, the return rate was 3% before the strike and 8% after, establishing a causal relationship between carrier delays and return costs. Through learning from the historical business database, the graph-based causal exploration model develops model parameters with the long-term marginal incremental cost of orders as the prediction target. The predicted long-term marginal incremental cost can include potential future additional storage fees, potential return / reshipment costs, additional customs clearance penalties, etc. The prediction results also include risk indicators such as delay / abnormality probabilities. Furthermore, the key driving factors affecting forecasting are returned in vectorized form, yielding implicit driving vectors whose components reflect the relative weights or contributions of different influencing factors. Further, based on the operable attributes of the influencing factors, such as carrier replacement, packaging enhancement, and warehouse switching, factors falling within the scope of operational adjustments are compiled into a list of controllable factors for operational optimization. For example, based on historical business databases, if the delay rate of a certain F1101 carrier is lower than that of the current carrier, then the list of controllable factors includes replacing the carrier with F1101.
[0067] S104. Based on the cost-event knowledge graph and implicit driving vectors, the cost of each order is allocated according to the actual event activities to obtain the cost waterfall for each order.
[0068] The cost waterfall is a breakdown of costs from purchase price to final net cost, showing the expenses at each stage, their sources, and the basis for allocation. Using business event nodes recorded in the cost-event knowledge graph as activity boundaries, costs are first aggregated, such as procurement, warehousing, storage, picking, packaging, transportation, customs clearance, overseas warehouse operations, last-mile delivery, returns processing, platform settlement, and marketing attribution. Then, based on the relationship and attributes between business events and business entities, such as volume * day, weight, transportation distance, consolidation rate, and number of processing times, the total cost of the activity pool is allocated to specific orders or SKUs according to contribution. The advantage of Activity-Based Cost Allocation (ABC) is that it is based on activities that actually contribute to the cost, rather than simply allocating costs according to the proportion of amount or number of items, thus more accurately reflecting the resource consumption and external risk costs borne by each order. Specifically, a cost-event knowledge graph is used to segment the end-to-end event sequence corresponding to an order into activity pools. For example, if a transportation event is delayed in an overseas warehouse for several days, the driving value of the warehousing activity pool will increase significantly. If a transportation event uses a specific carrier and involves multiple transfers, the driving force of the transportation activity pool should include more complex factors such as distance, number of transfers, and historical delay risks. Furthermore, a driving force formula is defined and a driving value is calculated for each activity pool of each order. Further, the base cost of a particular event in an order is allocated proportionally based on the total cost of the activity pool and the driving value. Optionally, implicit driving vectors are used as adjustment coefficients to calibrate the allocation weight of each activity pool across different orders, which can amplify or reduce the association between high-risk or high-cost items identified in the graph and the order. Finally, the costs of each activity item are chained together to form a cost waterfall, and the edge of the cost source and allocation basis are listed after each item to ensure the traceability of the allocation.
[0069] S105. Based on macro information, the cost waterfall is corrected, and management decisions are optimized based on the corrected cost waterfall to obtain a business management plan; macro information includes, but is not limited to, tariff policies, promotional calendars, and holidays; business management plan includes executable scheduling of order merging, logistics route selection, warehouse allocation, and adjustments to pricing and marketing campaigns.
[0070] By leveraging order-level cost waterfalls to adjust scenarios based on macroeconomic conditions and short-term forecasts, and executing actionable optimization decisions based on the revised cost view, the analysis results are transformed into actionable operational strategies. For illustrative purposes, macroeconomic information includes, but is not limited to, changes in tariffs and customs policies, such as adjustments to tariff rates in destination countries, changes in low-value duty-free policies, and temporary handling fees; promotional calendars, such as peak shipping times and return rates during major promotions; statutory holidays and peak transportation periods; and market-level fluctuations in freight rates and exchange rates.
[0071] Specifically, policy data is parsed into executable rules (policy engine), and the tax impact of each order under various policy scenarios is calculated based on its HS (Harmonized System) code and value. Simultaneously, a short-term forecasting model is constructed using historical sequences and event signals to estimate order demand peaks and troughs, carrier freight rates, and return rates over a future period. The forecast results are used to correct the cost waterfall. For example, if it is predicted that a destination country will increase processing fees, the customs clearance cost item for orders from that destination country is increased accordingly; if it is predicted that the return rate will increase during a promotional period, the expected return cost item for orders during that period is increased. Based on the corrected cost waterfall and the identified list of controllable factors, an optimization objective is constructed that prioritizes cost minimization while constraining delivery timeliness and service level. Appropriate real-time or rolling window optimization algorithms are used to solve executable strategies, such as order consolidation decisions (whether to postpone or consolidate shipments to reduce unit cost), route or carrier switching to balance cost and timeliness, warehouse and inventory allocation to reduce cross-border transshipment costs, and pricing and marketing adjustments to recover costs or avoid low-profit promotions. Furthermore, the optimization results are issued as specific instructions or suggestions to the TMS, WMS, procurement system, and advertising system for execution, and the execution results and actual costs are fed back to the business database and knowledge graph for model retraining and causal estimation updates, thus achieving closed-loop adaptive improvement.
[0072] The aforementioned big data-based refined cost management method for cross-border e-commerce constructs a cost-event knowledge graph containing entity nodes, event nodes, relational edges, and expense edges. This graph structure allows for tracking which event generated each expense and which entity it belongs to, automatically identifying cost components and expense paths, achieving structured cost breakdown, enhancing cost attribution capabilities, and improving the accuracy of cost control. Implicit driving vectors and a list of controllable factors can identify the most cost-sensitive business events or entities from historical business behavior, providing direction for subsequent optimization. The cost waterfall structure breaks down the cost of each order into multiple event levels, facilitating analysis of which events lead to high costs, achieving full-link visualization and layer-by-layer backtracking of order costs. Based on macro-information, the cost waterfall is dynamically adjusted to reflect changes in policies, activities, and holidays, adapting to the volatile nature of cross-border e-commerce policies. Decision recommendations are generated based on the latest cost structure, improving logistics utilization and warehouse scheduling efficiency, increasing business operational profitability, reducing overall costs, and minimizing decision-making lag.
[0073] In one embodiment, multi-source data is preprocessed to obtain a business database, including:
[0074] S11. Extract the business events corresponding to each order from the multi-source data using the order identifier as a clue, and standardize the extracted business events by time zone and timestamp to obtain the time-normalized order event stream.
[0075] This illustrative approach, using order identifiers as a guide to extract business events corresponding to each order, involves retrieving events related to the same order from the platform's order table, WMS inbound / outbound logs, TMS waybill event streams, customs clearance records, financial system entries, and marketing platform attribution records. This is done based on order number, order-waybill mapping tables, or composite primary keys (such as order number, SKU, and time window) to form a time-sorted event chain. The process also involves using the order identifier as the primary key to perform scanning, joining, and window aggregation. For one-to-many relationships, such as one waybill corresponding to multiple orders, an external mapping table or business rule-based mapping logic is used to assign events to the corresponding orders, while maintaining the event number association. Time normalization can be achieved through a unified time service or time database. If necessary, missing or abnormal timestamps are cleaned or corrected retrospectively, and the time correction strategy is written into the business database as part of data governance. The time-normalized order event stream includes the order identifier, event type, event identifier, unified timestamp, original source identifier, and basic fields for each event.
[0076] S12. Based on the preset entity matching rules, perform entity parsing and entity deduplication on the order event stream using a fuzzy matching algorithm to obtain aligned entity identifiers; entities include, but are not limited to, orders, suppliers, and carriers.
[0077] Entities include orders, suppliers, carriers, warehouses, SKUs, etc. The same supplier may appear with different spellings or codes in the procurement system and invoice database, and the abbreviation and official name of the same carrier may also differ across systems. To address this, a pre-defined set of entity matching rules is used. This set includes both deterministic rules, such as unified merchant tax identification numbers, unified carrier codes, and SKU coding standards, and fuzzy matching rules, such as string similarity thresholds, alias tables, and normalized name comparisons. Specifically, a blocking strategy is used to restrict candidate entity pairs to a small set with a high probability of association, reducing matching complexity. Fuzzy matching algorithms such as Levenshtein distance, Jaro-Winkler, or semantic similarity metrics are applied to candidate pairs, along with rule weights for evaluation. Business context is considered, such as transactions within the same time window, identical bank accounts, or addresses, for matching and scoring. The score, combined with a manual threshold or a learning model, determines whether entities should be merged. Optionally, for candidate pairs that cannot be automatically decided, manual review is triggered, and the manual results are written back into rules or training data to improve subsequent matching. This involves assigning a standardized entity identifier to each original record and writing that identifier into the order event stream and the business database.
[0078] S13. Based on foreign exchange data from multiple sources, convert the cost data in the order event stream that is priced in different currencies according to the exchange rate on the cost settlement date to obtain a uniformly priced cost amount.
[0079] As an illustration, the conversion rules prioritize the settlement date; if there is no settlement date, the invoice date or the event date is used. Rounding strategies and decimal place retention rules should be considered during the conversion process, and the exchange rate value, source, and timestamp used for each conversion should be recorded to ensure traceability. In the business database, the original currency and original amount are retained for each original amount field, and a new uniformly priced amount field is added.
[0080] S14. Based on the product text and attribute information in the SKU data, a pre-trained HS encoding annotation model is used to generate multiple corresponding HS encoding candidates and confidence scores for each SKU.
[0081] HS codes can be determined based on information such as product description, attributes, uses, and country of origin. A text classification or sequence labeling model trained on a large-scale labeled sample is used. The model inputs the SKU's title, specifications, and structured attributes, and the output is the probability distribution or candidate set of HS codes. Illustratively, a Transformer-based text classifier or multimodal model can be used. During the training phase, the model utilizes historical inventory interpretation data and manually labeled data for supervised learning, and during inference, it returns the top-N HS code candidates and their corresponding confidence scores. To balance recall and accuracy, the system should retain multiple candidates and record the source and confidence score of each candidate in the business database.
[0082] S15. Based on policy data and warehousing data, semantic annotation is performed on each order event stream to obtain semantic tags.
[0083] In a schematic representation, semantic labeling includes determining whether an order is shipped from an overseas warehouse, whether it may be eligible for tax refunds or returns, whether it is a promotional order, whether it is shipped during a peak period, and whether it involves special regulations. This labeling is based on policy data and warehousing data. Specifically, policy data is structured to perform strategy matching on fields such as HS candidate, value, and country of destination for orders. Based on the inbound / outbound and inventory attributes of the warehousing system, the system determines whether the product was shipped from an overseas warehouse or passed through a bonded warehouse. Based on the judgments of the above rules or models, the system adds semantic tags to the order event flow and records the labeling basis and confidence level. For ambiguous cases, the system can also write the confidence level of the semantic tags into the database and trigger manual review when necessary.
[0084] S16. Supplement each order event stream with fields consisting of entity identifier, uniformly priced fee amount, HS code candidate and its confidence level and semantic label to obtain the business database; the business database includes structured order business data.
[0085] As an illustration, the business database can use a relational database or a distributed data warehouse to store structured tables, where the tables include order masters, event-split transaction tables, standardized entity mapping tables, HS candidate codes and confidence levels, and semantic tags and bases. The structured order business data set contains each order record with an order event stream, aligned entity identifiers, uniform pricing amount, HS code candidates and confidence levels, semantic tags, and traceability information.
[0086] In one embodiment, such as Figure 2 As shown, based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph, obtaining implicit driving vectors and a list of controllable factors, including:
[0087] S201. Extract business event nodes and business entity nodes related to predefined volatile events from the cost-event knowledge graph.
[0088] Fluctuation-type events refer to exogenous shocks that cause significant changes in costs or service levels in time or space, such as sudden adjustments to the destination country's tariff policies, carrier strikes or service disruptions, sudden port congestion, order surges due to important promotional periods, or sharp increases in carrier rates. Illustratively, time windows are marked by external signals such as policy event databases, capacity event logs, or operational alarms, such as the period from the start of fluctuation to the stabilization period. The knowledge graph is sliced according to time windows, and active or affected business event nodes and connected business entity nodes within those time windows are extracted. Extraction strategies can employ connected component identification based on edge weight thresholds, event label-based filtering, and multi-hop expansion based on graph traversal to ensure coverage of both directly impacted nodes and related nodes that may be indirectly affected. The extracted subgraph retains timestamps, edge weights, and node attributes.
[0089] S202. Input the business event nodes and business entity nodes of the same volatile event into the graph-based causal exploration model to obtain the long-term marginal incremental cost and delay probability index of the order corresponding to the volatile event; the long-term marginal incremental cost includes, but is not limited to, potential return cost and potential warehousing cost; the graph-based causal exploration model is obtained by training causal inference on historical volatile event data.
[0090] Furthermore, graph structure learning and causal inference methods are used to quantitatively estimate the long-term marginal incremental cost of volatile events for each order or SKU, such as potential return costs, potential warehousing costs, potential reshipment costs, and related risk indicators such as delay probabilities. This transforms historically observed shocks into predictions of incremental costs for future orders in transit or not yet completed. Specifically, the training of the graph structure-based causal exploration model can be divided into two collaborative stages. For example, the first stage is representation learning based on graph neural networks (GNNs). By encoding the attribute vectors of nodes and edges as well as time slice graphs, a mapping function that can predict subsequent cost changes or event probabilities of nodes is learned. The architecture of the GNN can adopt a temporal graph neural network, such as temporal graph convolution or dynamic graph attention mechanism, to simultaneously encode spatial and temporal dependencies. The second stage is causal calibration. Natural experiments, such as known policy change time points, carrier anomalies, or other exogenous shocks, are used as identification sources. Differential-in-difference (DiD), instrumental variables (IV), or synthetic control-based methods are used to causally correct the correlation output of the GNN, thereby obtaining a long-term marginal incremental cost estimate that is closer to the causal effect. For example, for each evaluated order (o), the model output includes the predicted long-term marginal incremental cost and the probability of delay, along with a confidence interval or uncertainty measure.
[0091] S203. Based on the business entity nodes, the long-term marginal incremental cost is decomposed into an interpretable value to obtain the implicit driving vector; the implicit driving vector includes the weight value of each influencing factor.
[0092] Indicatively, model interpretability techniques and structured decomposition methods are used to decompose long-term marginal incremental costs into implicit driving vectors. Specifically, a local interpretation method based on GNNs is employed, utilizing attention weights based on edge / node importance or contribution scores based on input perturbations to measure the impact of each neighboring node or path on the target output. Optionally, posterior causal regression decomposition can also be used, i.e., after obtaining the causally corrected incremental cost, multiple regression or partial residual decomposition is performed on possible influencing factors to estimate the marginal contribution of each factor. Furthermore, to balance global and local interpretability, both global importance and order-level local interpretations—that is, the implicit driving vector of each order—can be output simultaneously, and the estimated value, confidence interval, and contribution description of each component can be recorded. For example, assignment methods similar to SHAP can be used to allocate model outputs, attention mechanisms can be used to directly read the weights within the graph model, or local causal effect estimation can be used to dissect specific paths.
[0093] S204. Based on the delay probability index, the influencing factors in the implicit driving vector are mapped to items that can be adjusted through operational actions, thus obtaining a list of controllable factors.
[0094] An illustrative mapping table of influencing factors and controllable actions can be predefined by domain experts or learned from historical intervention-effect samples. For example, historical operations involving changes to specific carriers and their subsequent cost changes can serve as evidence to verify the mapping relationship. For instance, based on delay probability indicators and implicit driving weights, several priority-ranked controllable factors are listed for each order, such as changing carriers to reduce delay risk, upgrading packaging to reduce damage rates and potential return costs, and changing from remote shipping to nearby overseas warehouses to reduce warehousing and cross-border freight costs. The expected cost-benefit of each controllable action is estimated, such as predicting the long-term marginal incremental cost reduction through carrier change. Furthermore, to support optimization solutions, constraints are defined for each controllable factor, such as the availability of carrier changes, material costs and implementation timelines for packaging specification adjustments, and the availability of overseas warehouse inventory. These constraints, along with the expected effects, are written into the controllable factor list.
[0095] In one embodiment, such as Figure 3 As shown, based on the cost-event knowledge graph and implicit driving vectors, costs are allocated to each order according to actual event activities, resulting in a cost waterfall for each order, including:
[0096] S301. Based on the business event nodes of the cost-event knowledge graph, divide the entire business process of each order into events to obtain the activity pool corresponding to each order. The activity pool includes procurement, warehousing, storage, picking, packaging, export transportation, customs clearance, overseas warehouse operations, last-mile delivery, return processing, platform settlement and marketing attribution.
[0097] Activity pools represent the collection of all costs incurred to complete a certain type of business activity, such as procurement, warehousing, storage, picking, packaging, export transportation, customs clearance, overseas warehouse operations, last-mile delivery, returns processing, platform settlement, and marketing attribution. Each activity pool may correspond to a single type of cost edge, such as a customs clearance fee edge, or it may be a combination of multiple cost edges using the same allocation criterion. For example, warehousing fees may involve multiple time periods and multiple edge weights, requiring consolidation by day / volume to arrive at the total cost of the warehousing pool. Specifically, the cost-event knowledge graph already displays the relationships between orders, events, and costs using nodes and edges. By expanding the event nodes connected to order nodes along the graph and grouping them according to event type, the allocation basis and attribution path can be made explicit. For example, a customs clearance fee edge can be directly connected to a group of customs declaration event nodes, thus clarifying which orders the cost should be allocated to. Furthermore, the multi-hop capability of the graph allows indirect impacts, such as delays caused by a carrier that increase storage days, to be mapped to the corresponding activity pool. This ensures that the cost pool covers both direct and indirect costs. For example, by traversing the cost-event knowledge graph, business events occurring for each order are aggregated into the corresponding activity pool according to their time sequence and event type. The edge weights of the cost edges directly associated with these events are then added to the initial cost of that activity pool. For shared costs across orders, such as transportation or consolidated customs clearance fees for multiple shipments, the set of shared edges and their original weights are recorded at this stage. The activity pool's allocability and allocation criteria are then marked, such as allocation by weight, volume, or number of items.
[0098] S302. Determine the adjustment value of the driver for the corresponding activity pool based on the weight value of the implicit driving vector; the driver is determined by the cost edge between the business entity node and the business event node.
[0099] A driver is a metric or function used to allocate the total cost of an activity pool among orders, such as volume multiplied by days, weight, shipping distance, consolidation rate, or number of processing times. The initial definition of a driver is usually determined by the semantics of the cost edge and the activity type. For example, storage fees are driven by volume occupied multiplied by storage days, and transportation fees by actual volume and weight multiplied by distance. The driver values are derived from the node / edge attributes of the cost-event knowledge graph and the order event flow; for example, warehouse inbound and outbound times are used to calculate the number of days occupied, and waybill route attributes are used to calculate the distance. Furthermore, relying solely on static drivers can underestimate or overestimate the true cost borne by certain orders when implicit risks exist. For example, if two items of the same volume have a significantly increased probability of return due to carrier delays on their route, the volume-based average allocation method cannot reflect the potential future additional costs. Therefore, the obtained implicit driver vectors are mapped to the drivers of the corresponding activity pools, and the driver adjustment value for each order under each activity pool is calculated, i.e., the driver scaling factor. Specifically, a basic driver formula is defined for each activity pool, and the required attribute fields for the driver are marked in the knowledge graph, such as volume, weight, inbound date, outbound date, transportation distance, declared value, HS-related risk indicators, etc. Based on the implicit driver vector, an implicit factor and driver mapping table is established. This table indicates which components in the implicit driver vector, such as carrier performance factors, HS review factors, and promotional return factors, should affect which activity pool drivers, as well as the direction and intensity of the influence.
[0100] S303. Based on the activity pool and drivers and their corresponding drive values and adjustment values, the cost waterfall for each order is obtained.
[0101] The cost waterfall can be obtained using the following formula:
[0102]
[0103] in, The total cost of the cost waterfall; For the activity pool collection, For the first Total cost of multiple orders across multiple activity pools; For orders In the drive The driving value below; For all orders in the first Total drive value under the class activity pool driver; For the first Each activity pool for orders The adjustment value.
[0104] Illustratively, the cost waterfall provides a breakdown view from purchase price to final gross profit, enabling management to clearly identify where each expense comes from, why it is allocated to that order, and to take appropriate strategies accordingly. Within each activity pool... Total cost above , The result is obtained by aggregating the corresponding cost edge weights in the cost-event knowledge graph according to the time window. Next, the sum of the calibrated driving values of all orders within this activity pool is calculated. For each order, the cost of the activity pool is allocated to the order level according to the formula.
[0105] In one embodiment, the method further includes:
[0106] S21. In response to monitoring and identifying an order return, refund, or tax refund event, based on the cost-event knowledge graph, trace along the association edges of the business entity nodes corresponding to the order to determine the relevant business event nodes and cost edges.
[0107] To illustrate, monitoring and identifying return, refund, or tax refund events can be achieved by analyzing the business database and external receipts to generate return / refund / tax refund events and write them into the order event stream. Specifically, using key fields in the event, such as the original order number, returned product SKU, returned quantity, return time, return reason, and refund or tax refund amount, as retrieval clues, the business event nodes related to the order and the cost edges connected to these nodes are located in the cost-event knowledge graph. These cost edge weights are the previously recorded cost edge weights, such as freight edge weights, customs declaration edge weights, warehousing fee edge weights, packaging fee edge weights, and platform handling fee edge weights. Backtracking can employ a graph traversal method for multi-hop searches along established connections. The search depth and rules are configurable. For example, it prioritizes searching directly connected event nodes and expands to two- or three-hop connections according to rules to cover indirect costs. During backtracking, the original amount of each located cost edge is extracted, along with its allocation status, such as whether it has been allocated by order, the allocation ratio and time window, and the allocation destination, such as which orders and activity pools it has been assigned to. Optionally, the backtracking results are output in a structured list format, specifying all relevant business event node IDs, cost edge IDs, original cost amounts, the values currently allocated to orders, and the allocation basis.
[0108] Optionally, the retrospective identification should distinguish between different return / refund / tax refund scenarios, including but not limited to full order returns, partial returns (i.e., returning several items of a certain SKU), freight recovery (i.e., the carrier or buyer bears the return shipping cost), platform deduction (i.e., the platform deducts the commission that has been returned due to the return), and customs tax refund (i.e., the customs or tax authorities refund the customs duties / value-added tax that have been paid).
[0109] S22. Based on the traced business event nodes and cost edges, reverse the costs previously allocated to the activity pool of the order according to the preset reversal rules, and redistribute the remaining costs after reversal to related orders to obtain the updated cost waterfall for related orders.
[0110] Returns or tax refunds alter the beneficiary or sharer of costs. For example, if a combined shipping cost is returned, the portion originally allocated to the returned order should be refunded or borne by other non-returned orders. Directly deleting the cost without reallocation might result in insufficient cost sharing for the remaining orders or inconsistencies in cost accounting. Specifically, for each traced cost edge, the allocation amount currently assigned to the returned order should be determined. The basic principle of reversal is to withdraw the allocated amount from the cost waterfall of the returned order, obtaining the cost difference before and after reversal. For full-order returns, the reversal amount is usually equal to the allocated amount; for partial returns, it should be calculated based on the returned quantity or value. In cases where return handling fees or return shipping costs are borne by the buyer, the reversal amount should deduct the recoverable portion or add the cost adjustment item to be borne by the customer. Furthermore, if the activity pool corresponding to the cost edge still needs to bear the remaining cost, the remaining portion should be reallocated among other related orders in that activity pool. Pre-defined reversal rules typically specify the basis for redistribution, such as allocation based on the proportion of drive values after driver calibration in the activity pool, or allocation based on the value / weight / number of items of the remaining valid orders. If the cost edge has already been allocated to multiple orders according to the consolidated order logic, such as a waybill with multiple orders in one shipment, the consolidation relationship needs to be identified during reversal. First, the share of returned orders should be withdrawn according to the allocation proportion within the consolidation, and then the remaining part of the bill should be redistributed to the non-returned orders within the consolidation or related orders outside the consolidation according to the above redistribution rules.
[0111] In one embodiment, the cost waterfall is revised based on macro-level information, and management decisions are optimized based on the revised cost waterfall to obtain a business management solution, including:
[0112] S31. Perform structured parsing of the cross-border commodity tariff policy text in the policy data to obtain a set of executable rules; the set of executable rules includes the target country, commodity HS code, value range, and tax calculation rules.
[0113] In a schematic way, policy documents are broken down into items through text extraction or manual input. Each item includes the target country / region, HS code, value conditions, rate or formula, and exceptions. Furthermore, the items are mapped to a unified rule model to execute a set of rules. The rule model contains fields such as country, commodity HS code, value range, and tax calculation rules.
[0114] S32. Based on historical customs clearance data and commodity HS codes, perform confidence verification on HS code candidates to determine the HS code corresponding to each order.
[0115] HS codes directly determine tax rates, eligibility for tax exemption thresholds, or specific regulations. Incorrect codes can lead to inaccurate tax estimates, customs clearance delays, or compliance risks. Using historical customs clearance interpretations can combine model candidates with past declaration results, improving the accuracy and confidence of the final determination. Specifically, candidate filtering is performed based on rules and data matching. The top-N candidate HS codes are weighted using evidence from the order's product description, attributes, supplier, and historical customs clearance interpretations for similar SKUs. Furthermore, a confidence merging strategy is applied, such as merging the model confidence with the empirical confidence obtained from historical matching using a weighted average or Bayesian update method to calculate the final HS determination and its confidence threshold. Optionally, if the final confidence is higher than a preset threshold, it is written into the business database as the order's final HS; if it is lower than the threshold, it is marked as requiring manual review and the manual review path is maintained during subsequent customs clearance.
[0116] S33. Based on the set of executable rules, calculate the tax and fees according to the HS code and order value of each order.
[0117] Based on the set of enforceable rules and the determined order HS code, order value, and country of destination, the system automatically calculates the taxes and fees payable for each order under the current rules, including but not limited to customs duties, import VAT, processing fees, and other customs-related charges.
[0118] S34. Based on the promotional calendar and holidays, and according to historical order sequences and historical carrier rates, a pre-trained combined time series prediction model is used to make predictions to obtain order demand, logistics freight rates and return rates within the preset prediction period.
[0119] Combined time-series forecasting models can be a combination of baseline seasonality models and deep learning time-series models, or a hybrid of statistical models and causal regression, balancing interpretability and nonlinear fitting capabilities. They also readily incorporate exogenous features such as promotions, holidays, and freight rates as dummy variables or regression terms, thereby improving forecast stability and scenario adaptability. Specifically, a dataset is constructed containing features such as historical order volume, past promotional tags, holiday labels, carrier rate time series, historical return rates, and weather / logistics event tags. Furthermore, a multi-model parallel strategy is employed: a time-series forecasting model based on LSTM / Transformer can be used for short-term windows to capture complex nonlinear patterns, while models such as Prophet or SARIMAX can be used for medium- and long-term forecasts to capture trends and seasonality. Simultaneously, causal regression terms are introduced for promotional or policy events to estimate the immediate impact of the events. The model output should include point estimates and uncertainty ranges, and provide confidence levels and forecasting hypotheses for each forecast term.
[0120] S35. Update the cost waterfall based on taxes, logistics costs, and return rates to obtain the updated cost waterfall.
[0121] By accruing future potential taxes / returns into costs, the transportation cost item is adjusted so that the optimizer can use a more forward-looking cost baseline. Specifically, the cost waterfall is split into cost types, i.e., activity pool items, and corresponding adjustment rules are applied to each item. For taxes, the original item is directly replaced or supplemented with order-level tax estimates. For transportation, the expected future transportation costs are calculated by weighting or substituting the predicted freight rates against the unit prices of future periods. For returns-related items, the additional return processing costs and reshipment costs caused by future returns are estimated using the predicted return rate, and this estimated value is incorporated into the accrual item of the returns pool.
[0122] S36. Based on the revised cost waterfall and combined with order requirements, optimize management decisions to obtain a business management solution.
[0123] This example illustrates how the updated cost waterfall is combined with order demand forecasts and a list of controllable factors to form an executable business management solution. This transforms front-end analysis and forecasts into specific operational actions and instructions, ultimately achieving a balance between cost control and service constraints. Specifically, management decision optimization uses the revised cost view as the cost baseline, demand forecasts and the list of controllable factors as variables and constraints as inputs, and outputs suggestions or executable scheduling commands. For instance, the revised cost waterfall, order demand forecasts, and the list of controllable factors are used to construct a set of executable suggestions for decision support. This involves batch generating what-if scenarios, such as a consolidation window length of 24 / 48 hours, using carrier A or carrier B, prioritizing overseas warehouse delivery or direct mail, etc., and evaluating the cost, timeliness, and risk performance under each scenario to generate a priority set of solutions. Optionally, for routine scenarios that can be automatically issued, instructions can be issued to the execution system via API or messages. For high-risk scenarios or scenarios requiring manual judgment, detailed comparative reports and suggestions are generated for operational decision-making.
[0124] In one embodiment, a business management solution is obtained by optimizing management decisions based on the revised cost waterfall and order demand, including:
[0125] S41. Construct an optimization objective function that minimizes costs while satisfying time constraints and service level requirements based on the cost waterfall and order demand.
[0126] Specifically, the revised cost waterfall and order demand information are formalized into an objective function and constraint system that can be mathematically optimized and solved. The objective function takes the minimization of overall cost as the main line, while taking delivery time, customer service level, such as arrival window, return rate threshold, and other business constraints as constraints that must be met or have penalties.
[0127] Decision variables may include, but are not limited to, order consolidation decisions (whether to merge several pending orders into one shipment), route and carrier selection (assigning carriers and routes to each pending or consolidated shipment), warehouse allocation (which warehouse to ship from or whether to activate overseas warehouses), pricing / promotion adjustments (whether to adjust the online price of a product or suspend a promotion), and short-term reallocation of advertising budgets (reducing advertising spending on low-margin products). Inputs include the modified cost waterfall, order demand forecasts, a list of controllable factors, and resource constraints such as carrier capacity, warehouse outbound capacity, and contract requirements. The overall objective function is written as the weighted sum of costs for all orders and action combinations, illustratively.
[0128] in, Indicating in decision-making schemes Place an order Expected costs; This indicates the penalty costs incurred for breaches of timeliness or service standards; This indicates the direct costs and operational expenses incurred due to changing carriers / warehouses / packaging; This refers to the expected additional revenue or cost savings resulting from pricing or marketing adjustments.
[0129] The constraints include, but are not limited to, delivery time constraints, such as target arrival time windows and maximum acceptable transportation delays; resource availability constraints, such as the capacity of carriers / warehouses / personnel; service level constraints, such as priority delivery ratios by channel or customer category; compliance constraints, such as certain destination countries or HS goods requiring specific channels or additional inspections; and business rule constraints, such as minimum order consolidation thresholds and the prohibition of delayed shipments of promotional items.
[0130] S42. Based on the objective function, a greedy algorithm or mixed integer programming algorithm is used to solve the optimal solutions for order merging strategy, logistics route selection, warehouse allocation, carrier selection, pricing and marketing placement adjustment according to the list of controllable factors, so as to obtain the business management solution.
[0131] Indicatively, for global batch optimization within a rolling window, a mixed-integer programming (MIP) solver can be used with reasonable time limits and heuristic warm-start. For real-time single or small-batch order decisions, a greedy strategy or fast heuristic is employed to ensure low-latency response. Specifically, batch solving (rolling window) models the optimization problem as an MIP, with decision variables including binary variables for order consolidation, carrier allocation, warehouse outbound, and pricing adjustment. To accelerate the process, heuristic initial solutions can be used, such as warm-start based on historical best solutions or solutions obtained through greedy strategies, returning optimal or near-optimal solutions and lower / upper bound proofs within a finite time. Real-time / near-real-time solving employs greedy order consolidation, local search, or rule-driven decision-making. For critical risks, such as upcoming tax changes or sudden freight rate increases, scenario generation and multi-objective or robust optimization are used to make the solution resilient to uncertainty; alternatively, conservative thresholds can be used to set high-uncertainty orders as non-consolidation or prioritize rapid shipment to avoid spillover risks. If managers need to simultaneously improve service levels and reduce costs, they can use weighted objectives or Pareto frontier search to generate multiple alternative solutions and output their respective cost / time / KPI estimates for operations staff to choose from.
[0132] The solver outputs a business management solution including specific instructions for each pending order, a list of outbound tasks for each warehouse, order placement instructions from the carrier interface, pricing / deployment adjustment suggestions for certain SKUs, and corresponding API call parameters. Optionally, the output may also include KPI estimates, such as expected cost savings, expected timeliness ratios, solution confidence levels, and risk descriptions.
[0133] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0134] Based on the same inventive concept, this application also provides a big data-based cross-border e-commerce cost refinement management device for implementing the aforementioned big data-based cross-border e-commerce cost refinement management method. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more big data-based cross-border e-commerce cost refinement management device embodiments provided below can be found in the limitations of the big data-based cross-border e-commerce cost refinement management method described above, and will not be repeated here.
[0135] In one exemplary embodiment, such as Figure 4 As shown, a data-driven, refined cost management device for cross-border e-commerce is provided, comprising:
[0136] The business data module 401 is used to acquire multi-source data of cross-border e-commerce business and preprocess the multi-source data to obtain the business database; the multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing bills, foreign exchange data and policy data;
[0137] The knowledge graph module 402 is used to construct a cost-event knowledge graph based on the business database. The cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes.
[0138] The implicit cost-driven discovery module 403 is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph based on historical business data and a pre-trained graph-based causal exploration model to obtain implicit driving vectors and a list of controllable factors.
[0139] The cost analysis module 404 is used to allocate costs and expenses for each order based on actual event activities, using a cost-event knowledge graph and implicit driving vectors, to obtain a cost waterfall for each order.
[0140] The management optimization module 405 is used to correct the cost waterfall based on macro information and optimize management decisions based on the corrected cost waterfall to obtain a business management solution. Macro information includes, but is not limited to, tariff policies, promotional calendars and holidays. The business management solution includes executable scheduling of order merging, logistics route selection, warehouse allocation, and adjustments to pricing and marketing.
[0141] In one embodiment, the business data module 401 is further configured to:
[0142] The business events corresponding to each order are extracted from multi-source data using order identifiers as clues. The extracted business events are then standardized by time zone and timestamp to obtain a time-normalized order event stream.
[0143] Based on preset entity matching rules, the order event stream is parsed and deduplicated using a fuzzy matching algorithm to obtain aligned entity identifiers; entities include, but are not limited to, orders, suppliers, and carriers;
[0144] Based on foreign exchange data from multiple sources, the cost data in the order event stream, which are priced in different currencies, are converted into currency according to the exchange rate on the cost settlement date to obtain a uniformly priced cost amount.
[0145] Based on the product text and attribute information in the SKU data, a pre-trained HS coding annotation model is used to generate multiple corresponding HS coding candidates and confidence scores for each SKU.
[0146] Semantic labels are generated by semantically annotating each order event stream based on policy data and warehousing data.
[0147] Each order event stream is supplemented with fields consisting of entity identifier, uniformly priced fee amount, HS code candidate and its confidence level and semantic label to obtain the business database; the business database includes structured order business data.
[0148] In one embodiment, the implicit cost-driven discovery module 403 is further configured to:
[0149] Extract business event nodes and business entity nodes related to predefined volatile events from the cost-event knowledge graph;
[0150] By inputting business event nodes and business entity nodes of the same volatile event into a graph-based causal exploration model, the long-term marginal incremental cost and delay probability index of the order corresponding to the volatile event are obtained; the long-term marginal incremental cost includes, but is not limited to, potential return costs and potential warehousing costs; the graph-based causal exploration model is obtained by training causal inference on historical volatile event data;
[0151] Based on the business entity nodes, the long-term marginal incremental cost is decomposed into an interpretable value to obtain the implicit driving vector; the implicit driving vector includes the weight value of each influencing factor.
[0152] Based on the delay probability index, the influencing factors in the implicit driving vector are mapped to items that can be adjusted through operational actions, thus obtaining a list of controllable factors.
[0153] In one embodiment, the cost analysis module 404 is further configured to:
[0154] Based on the business event nodes of the cost-event knowledge graph, the entire business process of each order is divided into events to obtain the activity pool corresponding to each order. The activity pool includes procurement, warehousing, storage, picking, packaging, export transportation, customs clearance, overseas warehouse operations, last-mile delivery, return processing, platform settlement and marketing attribution.
[0155] The adjustment value of the driver for the corresponding activity pool is determined based on the weight value of the implicit driving vector; the driver is determined by the cost edge between the business entity node and the business event node.
[0156] Based on the activity pool and drivers, along with their corresponding drive and adjustment values, the cost waterfall for each order is obtained.
[0157] In one embodiment, an event cost correction module is also included, for:
[0158] In response to the detection and identification of an order return, refund, or tax refund event, the relevant business event nodes and cost edges are determined by tracing along the association edges of the business entity nodes corresponding to the order based on the cost-event knowledge graph.
[0159] Based on the traced business event nodes and cost edges, the costs previously allocated to the activity pool of the order are reversed according to the preset reversal rules, and the remaining costs after reversal are redistributed to related orders to obtain the updated cost waterfall for related orders.
[0160] In one embodiment, the management optimization module 405 is further configured to:
[0161] The cross-border commodity tariff policy text in the policy data is structured and parsed to obtain a set of executable rules; the set of executable rules includes the target country, commodity HS code, value range, and tax calculation rules;
[0162] Based on historical customs clearance data and commodity HS codes, the confidence level of HS code candidates is verified to determine the HS code corresponding to each order.
[0163] Based on the set of executable rules, taxes and fees are calculated according to the HS code and order value of each order.
[0164] Based on promotional calendars and holidays, and according to historical order sequences and historical carrier rates, a pre-trained combined time series forecasting model is used to make predictions, and the order demand, logistics freight rates and return rates within the preset forecast period are obtained.
[0165] The cost waterfall is updated based on taxes, logistics costs, and return rates to obtain the updated cost waterfall.
[0166] Based on the revised cost waterfall and combined with order requirements, management decisions are optimized to obtain a business management solution.
[0167] In one embodiment, a management optimization solution module is also included, for:
[0168] Construct an optimization objective function that minimizes costs while satisfying time constraints and service level requirements based on the cost waterfall and order demand.
[0169] Based on the objective function, a greedy algorithm or mixed integer programming algorithm is used to solve the optimal solutions for order merging strategy, logistics route selection, warehouse allocation, carrier selection, pricing and marketing adjustment according to the list of controllable factors, so as to obtain the business management solution.
[0170] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.
[0171] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0172] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0173] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A method for refined cost management in cross-border e-commerce based on big data, characterized in that: The method includes: Acquire multi-source data from cross-border e-commerce business and preprocess the multi-source data to obtain a business database; the multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing invoices, foreign exchange data, and policy data; A cost-event knowledge graph is constructed based on the business database; the cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes. Based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph to obtain implicit driving vectors and a list of controllable factors. Based on the cost-event knowledge graph and the implicit driving vector, the cost of each order is allocated according to the actual event activities, resulting in a cost waterfall for each order. The cost waterfall is corrected based on macroeconomic information, and management decisions are optimized based on the corrected cost waterfall to obtain a business management solution. The macroeconomic information includes, but is not limited to, tariff policies, promotional calendars, and holidays. The business management solution includes executable order merging, logistics route selection, warehouse allocation, and adjustments to pricing and marketing campaigns.
2. The method according to claim 1, characterized in that, The preprocessing of the multi-source data to obtain the business database includes: The business events corresponding to each order are extracted from the multi-source data using the order identifier as a clue, and the extracted business events are standardized by time zone and timestamp to obtain a time-normalized order event stream. Based on preset entity matching rules, the order event stream is parsed and deduplicated using a fuzzy matching algorithm to obtain aligned entity identifiers; the entities include, but are not limited to, orders, suppliers, and carriers; Based on the foreign exchange data in the multi-source data, the cost data in the order event stream that are priced in different currencies are converted into currency according to the exchange rate on the cost settlement date to obtain a uniformly priced cost amount. Based on the product text and attribute information in the SKU data, a pre-trained HS coding annotation model is used to generate multiple corresponding HS coding candidates and confidence scores for each SKU. Semantic annotation is performed on each of the order event streams based on the policy data and warehousing data to obtain semantic tags; Each order event stream is supplemented with fields consisting of the entity identifier, uniformly priced fee amount, HS code candidate and its confidence level, and semantic label to obtain the business database; the business database includes structured order business data.
3. The method according to claim 2, characterized in that, Based on historical business data, a pre-trained graph-based causal exploration model is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph to obtain implicit driving vectors and a list of controllable factors, including: Extract the business event nodes and business entity nodes related to predefined volatile events from the cost-event knowledge graph; The business event nodes and business entity nodes of the same volatile event are input into the graph-based causal exploration model to obtain the long-term marginal incremental cost and delay probability index of the order corresponding to the volatile event; the long-term marginal incremental cost includes, but is not limited to, potential return cost and potential warehousing cost; the graph-based causal exploration model is obtained by training causal inference on historical volatile event data; Based on the business entity nodes, the long-term marginal incremental cost is decomposed into an interpretable form to obtain the implicit driving vector; the implicit driving vector includes the weight value of each influencing factor. Based on the aforementioned delay probability index, the influencing factors in the implicit driving vector are mapped to items that can be adjusted through operational actions, thereby obtaining the list of controllable factors.
4. The method according to claim 3, characterized in that, Based on the cost-event knowledge graph and the implicit driving vector, the cost of each order is allocated according to actual event activities to obtain a cost waterfall for each order, including: Based on the business event nodes of the cost-event knowledge graph, the entire business process of each order is divided into events to obtain the activity pool corresponding to each order; the activity pool includes procurement, warehousing, storage, picking, packaging, export transportation, customs clearance, overseas warehouse operations, last-mile delivery, return processing, platform settlement and marketing attribution; The adjustment value of the driver corresponding to the activity pool is determined based on the weight value of the implicit driving vector; the driver is determined by the cost edge between the business entity node and the business event node; Based on the activity pool and the driver and its corresponding driver value and adjustment value, the cost waterfall for each order is obtained; The cost waterfall is obtained using the following formula: in, The total cost of the cost waterfall; For the activity pool collection, For the first Total cost of multiple orders across multiple activity pools; For orders In the drive The driving value below; For all orders in the first Total drive value under the class activity pool driver; For the first Each activity pool for orders The adjustment value.
5. The method according to claim 4, characterized in that, The method further includes: In response to monitoring and identifying one of the following events: return, refund, or tax refund for the order, the relevant business event node and the cost edge are determined by tracing along the association edge of the business entity node corresponding to the order based on the cost-event knowledge graph. Based on the traced business event node and the cost edge, the cost previously allocated to the activity pool of the order is reversed according to the preset reversal rules, and the remaining cost after reversal is redistributed to the associated orders to obtain the updated cost waterfall for the associated orders.
6. The method according to claim 5, characterized in that, The process of revising the cost waterfall based on macro information and optimizing management decisions based on the revised cost waterfall to obtain a business management solution includes: The cross-border commodity tariff policy text in the policy data is structured and parsed to obtain a set of executable rules; the set of executable rules includes the target country, commodity HS code, value range, and tax calculation rules; Based on historical customs clearance data and the HS codes of the goods, the confidence level of the HS code candidates is verified to determine the HS code corresponding to each order; Based on the set of executable rules, taxes and fees are calculated according to the HS code and order value of each order. Based on the promotional calendar and holidays, and according to historical order sequences and historical carrier rates, a pre-trained combined time series prediction model is used to predict order demand, logistics freight rates and return rates within a preset prediction period. The cost waterfall is updated based on the taxes, logistics costs, and return rates to obtain the updated cost waterfall. Based on the revised cost waterfall and combined with the order requirements, management decisions are optimized to obtain the business management solution.
7. The method according to claim 6, characterized in that, The business management solution is obtained by optimizing management decisions based on the revised cost waterfall and the order demand, including: Based on the cost waterfall and the order demand, construct an optimization objective function that minimizes costs while satisfying time constraints and service level requirements; Based on the aforementioned objective function, a greedy algorithm or a mixed integer programming algorithm is used to solve for the optimal solutions for order merging strategy, logistics route selection, warehouse allocation, carrier selection, pricing, and marketing placement adjustment according to the controllable factor list, thereby obtaining the aforementioned business management solution.
8. A device for refined cost management in cross-border e-commerce based on big data, characterized in that, The device includes: The business data module is used to acquire multi-source data of cross-border e-commerce business and preprocess the multi-source data to obtain a business database; the multi-source data includes, but is not limited to, order data, SKU data, procurement data, warehousing data, transportation data, customs data, marketing bills, foreign exchange data and policy data; The knowledge graph module is used to construct a cost-event knowledge graph based on the business database; the cost-event knowledge graph includes business entity nodes, business event nodes, relationship edges between business entity nodes, and cost edges between business entity nodes and business event nodes. The implicit cost-driven discovery module is used to evaluate the long-term marginal increment of each incomplete order on the cost-event knowledge graph based on historical business data through a pre-trained graph-based causal exploration model to obtain implicit driving vectors and a list of controllable factors. The cost analysis module is used to allocate costs and expenses for each order based on actual event activities, based on the cost-event knowledge graph and the implicit driving vector, to obtain the cost waterfall for each order; The management optimization module is used to correct the cost waterfall based on macro information and optimize management decisions based on the corrected cost waterfall to obtain a business management solution; the macro information includes, but is not limited to, tariff policies, promotional calendars and holidays; the business management solution includes executable scheduling of order merging, logistics route selection, warehouse allocation, and adjustments to pricing and marketing campaigns.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.