Cross-border e-commerce adaptive replenishment method and system based on multi-scale causal perception

By constructing a multi-scale causal graph model, dynamically updating and coupling interactions across layers, the problem of inaccurate prediction of complex market demand in cross-border e-commerce inventory replenishment decisions is solved. This enables accurate identification of sudden trends and accurate adaptive generation of replenishment strategies, improving the adaptability and robustness of replenishment decisions.

CN121707686BActive Publication Date: 2026-06-02FUJIAN EASTWEST LIFEWIT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUJIAN EASTWEST LIFEWIT TECH CO LTD
Filing Date
2026-02-12
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing cross-border e-commerce inventory replenishment decision-making methods are unable to effectively analyze the internal driving relationships and causal structures of complex and non-stationary market demand, resulting in forecast results that are easily affected by noise or have a delayed response. The replenishment decisions are insufficient in terms of agility, accuracy and robustness.

Method used

A multi-scale causal graph model is constructed. By receiving multi-source heterogeneous real-time data streams, time alignment and feature fusion are performed to generate multi-dimensional temporal feature vectors. Long-term, medium-term and short-term causal graphs are constructed in parallel, dynamically updated and coupled across layers to perform early detection and lifecycle prediction of potential network trends. Combined with causal edge credibility assessment and reinforcement learning decision modules, accurate replenishment decision strategies are generated.

Benefits of technology

It enhances the adaptability and robustness of replenishment decisions to sudden market demands, enables accurate identification and dynamic calibration in complex market environments, and improves the accuracy and agility of replenishment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707686B_ABST
    Figure CN121707686B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-scale causal perception's cross-border e-commerce adaptive replenishment method and system, by receiving multi-source heterogeneous real-time data stream, time alignment and feature fusion are generated multidimensional time series feature vector, long-term, medium-term and short-term causal graph are constructed and dynamically updated in parallel;Based on short-term causal graph, early detection and life cycle prediction of potential network trend are carried out;The edge of short-term causal graph is carried out credibility evaluation and weighted adjustment;Based on long-term causal graph and medium-term causal graph, trend background calibration and operation interference filtering are carried out on trend prediction result, and calibrated trend prediction result is generated and replenishment decision strategy is selected accordingly;Through reinforcement learning decision module, replenishment action instruction is generated, instruction is executed and feedback data is collected, and feedback data is used to update and optimize causal graph and decision module online.The application improves the adaptability and robustness of replenishment decision to sudden market demand through the collaborative perception and dynamic calibration of multi-scale causal graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of e-commerce inventory management technology, and in particular to an adaptive replenishment method and system for cross-border e-commerce based on multi-scale causal perception. Background Technology

[0002] Inventory replenishment decisions in cross-border e-commerce heavily rely on accurate market demand forecasts. Due to the involvement of transnational supply chains, replenishment cycles are long and costs are high, placing extremely stringent demands on the foresight and stability of forecasts. Existing technologies generally employ time-series analysis methods based on multi-source data such as historical sales and inventory. These methods attempt to predict future sales and guide replenishment through weighted calculations and the introduction of seasonal adjustment factors (e.g., using weighted sales volume and inventory availability days to formulate shipping plans). These methods can, to some extent, cope with the regular cyclical fluctuations of the market.

[0003] However, with increasingly complex online consumer behavior, particularly driven by factors such as social media, market demand exhibits stronger characteristics of suddenness, non-stationarity, and multi-scale coupling. Short-term "blockbuster" trends, medium-term platform promotions, and long-term seasonal trends intertwine and influence each other, forming a complex dynamic system. Existing forecasting methods based on statistical correlation and fixed-rule models struggle to effectively analyze the intrinsic driving relationships and causal structures between market dynamics at different time scales. This makes it difficult for the system to effectively isolate sudden demand signals from background trends and operational disturbances and accurately assess their impact. Forecast results are susceptible to noise interference or delayed responses, resulting in significant deficiencies in the agility, accuracy, and robustness to uncertainty in subsequent replenishment decisions, thus hindering further improvements in the sophistication of inventory management. Summary of the Invention

[0004] In view of this, the purpose of this invention is to propose an adaptive replenishment method and system for cross-border e-commerce based on multi-scale causal perception. By constructing and coordinating the use of causal graph models at different time scales, it can accurately identify sudden trends in market dynamics and effectively remove multi-scale background interference, thus solving the problems of inaccurate prediction and poor decision robustness of existing replenishment decision-making methods when facing complex and non-stationary market demands.

[0005] To achieve the aforementioned technical objectives, in the first aspect, the technical solution adopted by the present invention is: a cross-border e-commerce adaptive replenishment method based on multi-scale causal perception, comprising:

[0006] Receive multi-source heterogeneous real-time data streams;

[0007] Time alignment and feature fusion are performed on multi-source heterogeneous real-time data streams to generate a unified multi-dimensional time-series feature vector;

[0008] Based on multidimensional temporal feature vectors, long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated. The long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs interact through cross-layer coupling.

[0009] Early detection of potential network trends is performed based on short-term causal graphs, and lifecycle prediction is initiated for the detected potential network trends to generate trend prediction results that include predicted demand curves and lifecycle stages.

[0010] The credibility of edges representing causal relationships in short-term causal graphs is evaluated to generate a first causal edge credibility score, and the strength of causal edges in short-term causal graphs is weighted and adjusted using the first causal edge credibility score.

[0011] Based on long-term and medium-term causal graphs, trend background calibration and operational interference filtering are performed on the storm surge forecast results to generate calibrated storm surge forecast results.

[0012] The corresponding replenishment decision strategy is selected based on the life cycle stage in the calibrated wind tide forecast results. The replenishment decision strategy includes one of the following: normal mode strategy, exploratory mode strategy, aggressive mode strategy, and cleanup mode strategy.

[0013] Based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, specific replenishment action instructions are generated through the reinforcement learning decision module.

[0014] Execute replenishment orders and collect actual sales and inventory feedback data;

[0015] Using actual sales and inventory feedback data, the parameters of the long-term, medium-term, and short-term causal graphs are updated online, and the strategy parameters of the reinforcement learning decision-making module are optimized based on feedback.

[0016] In some embodiments, based on multidimensional temporal feature vectors, long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated, including:

[0017] Data aggregation at different time scales is performed on multidimensional time series feature vectors to generate long-term aggregated feature vectors, medium-term aggregated feature vectors, and short-term aggregated feature vectors.

[0018] Based on long-term aggregated feature vectors, the causal structure between monthly or quarterly scale variables is discovered through the FCI algorithm, and a long-term causal graph is constructed.

[0019] Based on the mid-term aggregated feature vector and combined with the pre-set prior knowledge of operational activities, a mid-term causal graph representing the causal influence between weekly operational variables is constructed using the LiNGAM algorithm or a constraint-based causal discovery algorithm.

[0020] Based on short-term aggregated feature vectors, the instantaneous causal relationships between variables in high-dimensional streaming data at the minute or hour scale are discovered through the neural Granger causality algorithm, and a short-term causal graph is constructed.

[0021] In some embodiments, the long-term causal graph, the intermediate-term causal graph, and the short-term causal graph interact through cross-layer coupling, including:

[0022] Establish a cross-layer bidirectional coupling mechanism between long-term causal graphs, medium-term causal graphs, and short-term causal graphs;

[0023] In the cross-layer bidirectional coupling mechanism, the operation of transferring seasonal baseline prior constraints from the long-term causal graph to the medium-term and short-term causal graphs includes:

[0024] Extract causal pathways and their effect strengths that characterize stable seasonal patterns from long-term causal graphs;

[0025] Convert the effect intensity into a demand baseline component;

[0026] The baseline demand component is injected as a prior mean into the probability distribution parameters of the corresponding demand variables in the intermediate and short-term causal graphs;

[0027] In the cross-layer bidirectional coupling mechanism, the operation of feeding back anomalous causal strength signals from the intermediate-term causal graph and the short-term causal graph to the long-term causal graph includes:

[0028] Continuously monitor the deviation between the instantaneous strength of specific causal edges and historical baselines in medium-term and short-term causal graphs;

[0029] When the magnitude and duration of the deviation exceed a preset significance threshold, an abnormal causal strength signal is generated.

[0030] The causal strength anomaly signal and its associated contextual features are used as input events to trigger structural reassessment and parameter updates of the long-term causal graph.

[0031] In some embodiments, potential network trends are detected early based on short-term causal graphs, and lifecycle prediction is initiated for the detected potential network trends to generate trend prediction results that include predicted demand curves and lifecycle stages, including:

[0032] Extract time-series feature indicators representing social media volume, search popularity, and information dissemination topology from short-term causal graphs;

[0033] Calculate the rate of change and acceleration of change of time-series characteristic indicators, and compare them with dynamic detection thresholds set based on historical data distribution;

[0034] When the rate of change or acceleration of change of the time sequence characteristic index exceeds the dynamic detection threshold, it is determined that a potential network storm has been detected and recorded as a storm initiation event.

[0035] In response to the trend initiation event, a hybrid lifecycle prediction model is launched;

[0036] In the hybrid life cycle prediction model, for the initial stage of the surge, the Hawkes process model is used to fit recent time series characteristic data to predict the intensity of demand stimulation in the next few hours.

[0037] In the hybrid life cycle prediction model, for the middle stage of the epidemic, a variant of the SEIR infectious disease model is used to predict the scale of infection and peak time of the potential demand group based on the simulation of node state transitions in the information propagation network.

[0038] In the hybrid life cycle prediction model, for the decline period of a hot trend, an exponential decay model is adopted, which combines the decay parameters of similar historical hot trend patterns to predict the demand decline curve.

[0039] By integrating the forecast outputs of the Hawkes process model, the SEIR infectious disease model variant, and the exponential decay model, a continuous forecast demand curve covering the entire life cycle is generated.

[0040] Based on the morphological characteristics of the continuously predicted demand curve, the initiation period, the peak period, and the decline period of the trend are divided and identified, forming a trend prediction result that includes the predicted demand curve and the life cycle stage.

[0041] In some embodiments, the credibility of edges representing causal relationships in a short-term causal graph is evaluated to generate a first causal edge credibility score, including:

[0042] For each edge representing a causal relationship in the short-term causal graph, a Granger causality test is performed to generate a Granger causality test statistic representing the stability of the time lead-lag relationship.

[0043] For social media information propagation paths associated with edges representing causal relationships, we analyze the attribute authority of nodes in the propagation path and the topological diversity of the propagation path to generate a propagation path credibility metric.

[0044] For text content associated with edges representing causal relationships, natural language processing is used to analyze the consistency of text sentiment polarity and the relevance of text topics to the core attributes of products, generating a content semantic credibility metric.

[0045] Based on Granger causality test statistics, propagation path credibility measure and content semantic credibility measure, a first causal edge credibility score for the edge representing the causal relationship is generated through weighted fusion calculation.

[0046] The strength of causal edges in the short-term causal graph is weighted and adjusted using the first causal edge credibility score, including:

[0047] The original causal strength of each edge representing a causal relationship in the short-term causal graph is multiplied by the confidence score of the first causal edge corresponding to that edge to obtain the weighted adjusted causal edge strength.

[0048] Update the weight parameters of the corresponding edges in the short-term causal graph using the weighted adjusted causal edge strength.

[0049] In some embodiments, based on long-term and medium-term causality graphs, trend background calibration and operational interference filtering are performed on the storm surge forecast results to generate calibrated storm surge forecast results, including:

[0050] Extract the state of slow variable nodes representing seasonal trends and category lifecycles from the long-term causal graph, and use them as the trend background baseline;

[0051] Subtract the demand component corresponding to the trend background baseline from the continuous forecast demand curve of the wind tide forecast results to obtain the net wind tide demand curve after removing the influence of the long-term trend.

[0052] From the mid-term causal diagram, identify causal paths driven by operational activities that overlap in time with the lifecycle stages of the current trend forecast results;

[0053] Assess the strength of the impact of the causal path driven by operational activities on demand, and separate this strength from the net windfall demand curve to obtain the core windfall demand curve after removing operational disturbances;

[0054] By overlaying the trend background baseline with the influence intensity of the causal path driven by operational activities, the background and operational interference demand curves are reconstructed.

[0055] The core trend demand curve is fused with the reconstructed background and operational interference demand curves to generate a calibrated trend forecast result. The forecast demand curve in the calibrated trend forecast result has been calibrated for trend background and filtered for operational interference, and the life cycle stage identifier is inherited from the trend forecast result.

[0056] In some embodiments, selecting a corresponding replenishment decision strategy based on the lifecycle stage in the calibrated storm surge forecast results includes:

[0057] Identify the lifecycle stages marked in the calibrated storm surge forecast results. The lifecycle stages include the storm surge initiation period, the storm surge peak period, and the storm surge decline period.

[0058] When the life cycle stage is the trend initiation period, the replenishment decision strategy is to choose an exploration mode strategy that takes maximizing information value as the optimization goal and adopts a high exploration rate decision mechanism.

[0059] When the life cycle stage is the peak of the trend, an aggressive mode strategy that maximizes the demand satisfaction rate and allows temporary relaxation of inventory cost constraints is chosen as the replenishment decision strategy.

[0060] When the product lifecycle stage is in the decline phase, the replenishment decision strategy is to select a clearing mode strategy that maximizes the difference between sales revenue and inventory clearance costs and links the promotion execution system.

[0061] When the calibrated trend forecast results do not identify an effective lifecycle stage, the normal mode strategy with the optimization objective of minimizing the sum of inventory holding costs and stockout loss costs is selected as the replenishment decision strategy.

[0062] In some embodiments, based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, a reinforcement learning decision module generates specific replenishment action instructions, including:

[0063] Construct a state representation for the reinforcement learning decision module. The state representation includes the current inventory level, in-transit inventory information, the predicted demand curve and life cycle stage in the calibrated windstorm forecast results, and the strength of key causal edges in the weighted and adjusted short-term causal graph.

[0064] The selected replenishment decision strategy is mapped to the policy network parameters and reward function weights of the agent in the reinforcement learning decision module, wherein the reward function includes an economic benefit term and a causal consistency reward term.

[0065] Based on the state representation, the policy network parameterized by the policy network parameters is used to generate candidate replenishment quantity actions for the target product.

[0066] Input the candidate replenishment quantity actions, status representations, weighted adjusted short-term causal graphs, long-term causal graphs, and medium-term causal graphs into the causal simulator;

[0067] By using a causal simulator, the trajectory of changes in inventory, sales and key causal variables over multiple decision-making cycles in the future is simulated after executing candidate replenishment actions.

[0068] Based on the change trajectory, the economic benefit item and the causal consistency reward item are calculated. The causal consistency reward item is generated by measuring the degree of agreement between the actual relationship between variables in the change trajectory and the causal laws revealed by the long-term causal diagram, the medium-term causal diagram, and the short-term causal diagram.

[0069] By combining the economic benefits and the causal consistency reward, we obtain an immediate reward estimate for the candidate replenishment action.

[0070] Based on the instant reward estimation and policy network parameters, the policy network parameters are updated through the policy gradient method or the actor-critic algorithm, and the final replenishment action instruction is output, which includes the replenishment quantity and the replenishment time.

[0071] In some embodiments, parameters of the long-term causal graph, medium-term causal graph, and short-term causal graph are updated online using actual sales and inventory feedback data, and the policy parameters of the reinforcement learning decision module are optimized based on feedback, including:

[0072] The actual sales and inventory feedback data are compared with the predicted demand curve in the calibrated trend forecast results to generate a prediction error sequence;

[0073] Based on the prediction error sequence, online gradient descent is performed to update the weight parameters of relevant causal edges in the short-term causal graph to reduce the prediction error sequence.

[0074] Separate the systematic deviation components in the prediction error sequence that are related to long-term trends and operational activities;

[0075] Using the systematic bias component, Bayesian updates are performed on the node parameters representing the seasonal baseline in the long-term causal graph and the causal edge strength parameters representing the impact of operational activities in the medium-term causal graph, respectively.

[0076] The actual sales and inventory feedback data, the replenishment action instructions executed, and the updated long-term, medium-term, and short-term causal graphs together constitute the experience sample for the reinforcement learning decision-making module.

[0077] Store the experience samples into the experience replay buffer;

[0078] Sample batches of experience samples from the experience replay buffer and calculate the advantage function estimate for each state-action pair in the batch of experience samples.

[0079] Based on advantage function estimation, the policy network parameters and value network parameters of the agent in the reinforcement learning decision module are iteratively updated using the proximal policy optimization algorithm or the soft actor-critic algorithm to complete feedback optimization.

[0080] In a second aspect, the present invention also provides a cross-border e-commerce adaptive replenishment system based on multi-scale causal perception, applicable to the method of the first aspect. The system includes a data fusion module, a multi-scale causal graph construction and update module, a trend detection and prediction module, a causal signal credibility assessment module, a trend prediction calibration module, a strategy selection module, a reinforcement learning decision-making module, a decision execution and feedback module, and an online learning and optimization module. The data fusion module receives multi-source heterogeneous real-time data streams and performs time alignment and feature fusion on the multi-source heterogeneous real-time data streams to generate a unified multi-dimensional temporal feature vector. The multi-scale causal graph construction and update module constructs long-term, medium-term, and short-term causal graphs in parallel based on the multi-dimensional temporal feature vectors and dynamically updates the long-term, medium-term, and short-term causal graphs, wherein the long-term, medium-term, and short-term causal graphs interact through cross-layer coupling. The trend detection and prediction module performs early detection of potential network trends based on the short-term causal graph and initiates lifecycle prediction for the detected potential network trends, generating trend prediction results including predicted demand curves and lifecycle stages. Causal signal credibility assessment... The module is used to evaluate the credibility of edges representing causal relationships in the short-term causal graph, generate a first causal edge credibility score, and use the first causal edge credibility score to weight and adjust the strength of causal edges in the short-term causal graph. The trend prediction calibration module is used to perform trend background calibration and operational interference filtering on the trend prediction results based on the long-term and medium-term causal graphs, and generate calibrated trend prediction results. The strategy selection module is used to select the corresponding replenishment decision strategy based on the life cycle stage in the calibrated trend prediction results. The replenishment decision strategy includes one of the following: normal mode strategy, exploration mode strategy, aggressive mode strategy, and cleanup mode strategy. The reinforcement learning decision module is used to generate specific replenishment action instructions based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results. The decision execution and feedback module is used to execute the replenishment action instructions and collect actual sales and inventory feedback data. The online learning and optimization module is used to update the parameters of the long-term, medium-term, and short-term causal graphs online using actual sales and inventory feedback data, and to provide feedback optimization to the strategy parameters of the reinforcement learning decision module.

[0081] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: It receives multi-source heterogeneous real-time data streams, performs time alignment and feature fusion to generate multi-dimensional temporal feature vectors, and constructs and dynamically updates long-term, medium-term, and short-term causal graphs in parallel; it performs early detection and lifecycle prediction of potential network trends based on the short-term causal graph; it evaluates the credibility and adjusts the weights of the edges of the short-term causal graph; it performs trend background calibration and operational interference filtering on the trend prediction results based on the long-term and medium-term causal graphs, generates calibrated trend prediction results, and selects replenishment decision strategies accordingly; it generates replenishment action instructions through a reinforcement learning decision module, executes the instructions and collects feedback data, and uses the feedback data to update and optimize the causal graph and decision module online. The present invention improves the adaptability and robustness of replenishment decisions to sudden market demands through the collaborative perception and dynamic calibration of multi-scale causal graphs. Attached Figure Description

[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0083] Figure 1 This is a schematic diagram of steps S101 to S110 of the method described in the specific implementation embodiment;

[0084] Figure 2 This is a schematic diagram of steps S201 to S204 of the method described in the specific implementation embodiment;

[0085] Figure 3 This is a schematic diagram of the adaptive replenishment system described in the specific implementation.

[0086] The reference numerals for the above figures are as follows:

[0087] 1. Adaptive replenishment system;

[0088] 11. Data fusion module;

[0089] 12. Multi-scale cause-effect graph construction and update module;

[0090] 13. Wind and tide detection and forecasting module;

[0091] 14. Causal signal credibility assessment module;

[0092] 15. Wind and tide forecasting calibration module;

[0093] 16. Strategy Selection Module;

[0094] 17. Enhance learning decision-making module;

[0095] 18. Decision Execution and Feedback Module;

[0096] 19. Online Learning and Optimization Module. Detailed Implementation

[0097] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0098] Please see Figure 1 In a first aspect, this embodiment provides an adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception, including:

[0099] S101, Receive multi-source heterogeneous real-time data streams;

[0100] S102. Perform time alignment and feature fusion on multi-source heterogeneous real-time data streams to generate a unified multi-dimensional time-series feature vector;

[0101] S103. Based on multi-dimensional temporal feature vectors, long-term causal graphs, medium-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated. The long-term causal graphs, medium-term causal graphs, and short-term causal graphs interact through cross-layer coupling.

[0102] S104. Based on the short-term causal graph, perform early detection of potential network trends, and initiate lifecycle prediction for the detected potential network trends to generate trend prediction results that include the predicted demand curve and lifecycle stage.

[0103] S105. Evaluate the credibility of the edges representing causal relationships in the short-term causal graph, generate the credibility score of the first causal edge, and use the credibility score of the first causal edge to make a weighted adjustment of the strength of the causal edges in the short-term causal graph.

[0104] S106. Based on the long-term causal graph and the medium-term causal graph, perform trend background calibration and operational interference filtering on the storm surge forecast results to generate calibrated storm surge forecast results.

[0105] S107. Select the corresponding replenishment decision strategy based on the life cycle stage in the calibrated wind tide forecast results. The replenishment decision strategy includes one of the following: normal mode strategy, exploratory mode strategy, aggressive mode strategy, and cleanup mode strategy.

[0106] S108. Based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, specific replenishment action instructions are generated through the reinforcement learning decision module.

[0107] S109. Execute the replenishment action instruction and collect actual sales and inventory feedback data;

[0108] S110. Using actual sales and inventory feedback data, update the parameters of the long-term, medium-term, and short-term causal graphs online, and optimize the strategy parameters of the reinforcement learning decision module based on feedback.

[0109] In step S101, the multi-source heterogeneous real-time data stream encompasses internal operational data such as e-commerce platform sales orders, inventory levels, and product page traffic; external public opinion data such as social media topic discussion volume and sentiment trends; search engine keyword popularity index; and logistics transit status information. These data include structured records, semi-structured logs, and unstructured text in format, and vary in frequency from minutes to days.

[0110] In step S102, time alignment unifies data from different sources to a common time base, such as aggregating data at hourly intervals. Feature fusion then integrates the aligned multi-dimensional data, extracting feature indicators that comprehensively reflect the market state through methods such as normalization, principal component analysis, or embedding representation learning, and combining them to generate a unified multi-dimensional time-series feature vector. This vector represents the overall state of the market at each time point, providing a standardized input feature space for causal discovery.

[0111] In step S103, based on a unified multi-dimensional time-series feature vector, three causal graphs with different time scales are constructed in parallel. The long-term causal graph depicts stable causal patterns at the monthly or quarterly scale, the medium-term causal graph reflects the causal impacts related to operational activities at the weekly scale, and the short-term causal graph captures instantaneous dynamic correlations at the minute or hourly scale. Each causal graph is represented as a directed graph, with nodes corresponding to market feature variables and edges representing causal relationships and strengths; the structure and parameters of each causal graph are continuously adjusted with new data. Cross-layer coupling and interaction realize information transmission and constraints between causal graphs of different scales. For example, long-term trends can provide an explanatory baseline for short-term fluctuations, and short-term anomalies can trigger a reassessment of the medium- and long-term structure.

[0112] In step S104, the instantaneous dynamic relationships revealed by the short-term causal graph are used to monitor the strength changes of specific nodes or edges to detect potential network trends that may trigger sudden changes in demand. The detection criterion can be that the rate of change exceeds a dynamic threshold based on historical data distribution. For detected potential network trends, lifecycle prediction is initiated to estimate their entire time trajectory from rise to fall and their demand impact; the generated trend prediction results include a quantified predicted demand curve and a lifecycle stage identifying the current and future stages.

[0113] In step S105, the credibility assessment of edges representing causal relationships in the short-term causal graph can be based on multi-dimensional indicators such as the significance of statistical tests, the reliability of information propagation paths, or the semantic consistency of associated text content. The original strength of the causal edges is adjusted using the credibility score of the first causal edge, that is, the weight of high-credibility edges is enhanced, and the weight of low-credibility edges is weakened, thereby improving the reliability of the short-term causal graph signal.

[0114] In step S106, the long-term causal graph provides the trend background baseline, and the medium-term causal graph identifies demand interference components driven by operational activities. Trend background calibration and operational interference filtering remove these baseline and interference components from the trend forecast results, extracting the core demand forecast driven purely by short-term network trends. The calibrated trend forecast results more accurately reflect the potential impact of sudden trends themselves.

[0115] In step S107, among the replenishment decision strategies, the normal mode strategy is suitable for stable markets with the goal of cost optimization; the exploration mode strategy is suitable for the initial stage of a trend and focuses on information collection and risk control; the aggressive mode strategy is suitable for the peak stage of a trend and aims to maximize demand satisfaction; and the clearing mode strategy is suitable for the decline stage of a trend and aims to optimize inventory clearance revenue.

[0116] In step S108, the reinforcement learning decision module receives the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend forecast as input. By simulating the long-term benefits of different replenishment actions and integrating the causal constraints and demand forecasts in the input information, it generates specific replenishment action instructions. The instructions include the replenishment quantity and execution time.

[0117] In steps S109 to S110, replenishment instructions are sent to the execution system, and actual sales data and inventory change data generated after instruction execution are collected, constituting real-time feedback on the decision-making effect. The model parameters of the long-term, medium-term, and short-term causal graphs are updated online, which can be achieved by adjusting the causal edge strength using gradient descent or Bayesian methods. Simultaneously, this inventory feedback data is used to optimize the policy parameters of the reinforcement learning decision module, for example, by updating the policy network using policy gradient methods, enabling the system to continuously evolve based on market feedback.

[0118] This embodiment constructs a complete technical solution encompassing multi-source data fusion, multi-scale causal perception, trend detection and prediction, signal calibration and purification, adaptive strategy selection, and reinforcement learning decision-making, ultimately achieving closed-loop optimization through feedback. It achieves a collaborative understanding and decoupling of slow-changing market patterns and rapid-changing drivers through parallel construction and cross-layer coupling of the causal graph; enhances prediction robustness through credibility assessment and calibration of causal signals; and achieves precise adaptive generation of replenishment actions by dynamically matching trend lifecycle stages with replenishment strategies and utilizing reinforcement learning to integrate causal constraints for decision-making. This embodiment forms a closed loop of perception, decision-making, execution, and learning, effectively improving the accuracy, agility, and overall robustness of replenishment decisions in complex and non-stationary market environments.

[0119] Please see Figure 2 In some embodiments, based on multi-dimensional temporal feature vectors, long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated, including:

[0120] S201. Aggregate data at different time scales on the multi-dimensional time series feature vectors to generate long-term aggregated feature vectors, medium-term aggregated feature vectors and short-term aggregated feature vectors.

[0121] S202. Based on long-term aggregated feature vectors, the causal structure between monthly or quarterly scale variables is discovered through the FCI algorithm, and a long-term causal graph is constructed.

[0122] S203. Based on the mid-term aggregated feature vector and combined with the pre-set prior knowledge of operational activities, construct a mid-term causal graph representing the causal influence between weekly operational variables through the LiNGAM algorithm or a constraint-based causal discovery algorithm.

[0123] S204. Based on short-term aggregated feature vectors, the instantaneous causal relationship between variables in high-dimensional streaming data at the minute or hour scale is discovered by using the neural Granger causality algorithm, and a short-term causal graph is constructed.

[0124] In step S201, data aggregation at different time scales on the multidimensional time-series feature vector refers to statistically summarizing fine-grained time-series data according to predefined time windows. Long-term aggregation typically uses monthly or quarterly windows, calculating the mean, sum, or trend slope of each feature dimension within the window to generate a long-term aggregated feature vector. This vector reflects the macro trend across multiple short-term fluctuations. Medium-term aggregation uses weekly windows, performing similar operations to generate medium-term aggregated feature vectors, used to capture patterns related to business cycles or fixed operating rhythms. Short-term aggregation uses minute or hourly windows, possibly directly using raw high-frequency data or performing simple smoothing to generate short-term aggregated feature vectors to preserve details of real-time market dynamics. Through aggregation, the original high-dimensional time-series data is transformed into feature representations with corresponding time resolutions to meet the analytical needs of different causal discovery algorithms.

[0125] In step S202, the FCI algorithm is a causal discovery method based on conditional independence testing. It can infer the causal framework between variables from observed data and partially orient them, making it particularly suitable for scenarios where unobserved confounding factors may exist. The long-term aggregated feature vector is input into the FCI algorithm, which systematically tests the independence of different variable sets under a given condition set to discover the causal structure between various macroeconomic characteristic variables (such as category seasonality indices, macroeconomic indicators, etc.) at monthly or quarterly scales. Based on this discovered causal structure, initial effect strength estimates are assigned to the causal relationships between nodes, thus constructing a long-term causal graph. This causal graph characterizes the relatively stable driving relationships between slowly changing market factors.

[0126] In step S203, the pre-defined prior knowledge of operational activities includes information such as planned promotional activity schedules, advertising cycles, and platform major promotional holidays, typically existing in the form of a structured time event table. When constructing the intermediate causal graph, this prior knowledge is first encoded as additional binary features or time window identifiers, which are used as input along with the intermediate aggregated feature vector. The LiNGAM algorithm assumes that the data generation process is linear and the noise is non-Gaussian distributed, enabling it to identify a unique causal direction; while constraint-based causal discovery algorithms infer causal structures through more general conditional independence tests. Combining the prior knowledge of operational activities, some causal edges can be constrained or initialized (e.g., forcing the "promotional activity" node to point to the "sales" node), thereby guiding the algorithm to more accurately identify the causal influence network driven by operational actions on a weekly scale, forming the intermediate causal graph.

[0127] In step S204, the neural Granger causality algorithm utilizes neural networks (such as recurrent neural networks or temporal convolutional networks) to detect nonlinear Granger causal relationships between high-dimensional time series. In this embodiment, Granger causality is used to determine whether the historical information of a market feature variable contributes to predicting the future value of another feature variable beyond its own historical information. The neural Granger causality algorithm models the complex temporal dependencies between all variables by training a neural network, and based on the dependency patterns learned by the network, it infers which variables have immediate causal relationships by analyzing the connection weights within the network or imposing specific sparsity constraints. The resulting short-term causal graph reveals the dynamic causal network between various micro-indicators in a rapidly changing market environment.

[0128] This embodiment employs differentiated data aggregation strategies and causal discovery algorithms tailored to different time scales, enabling targeted causal modeling of macroeconomic trends, meso-level operational interventions, and micro-level real-time dynamics. Using the FCI algorithm to process long-term sparse data helps uncover robust macroeconomic causal frameworks; combining prior knowledge with mid-term causal discovery improves the accuracy of operational attribution; and applying the neuro-Granger causal algorithm effectively captures complex instantaneous causal relationships in high-dimensional streaming data. This differentiated technology selection allows the constructed long-term, mid-term, and short-term causal graphs to accurately characterize market operation mechanisms at their respective time scales, providing a reliable and interpretable causal foundation for subsequent cross-layer coupling, trend detection, and decision calibration.

[0129] In some embodiments, the long-term causal graph, the intermediate-term causal graph, and the short-term causal graph interact through cross-layer coupling, including:

[0130] Establish a cross-layer bidirectional coupling mechanism between long-term causal graphs, medium-term causal graphs, and short-term causal graphs;

[0131] In the cross-layer bidirectional coupling mechanism, the operation of transferring seasonal baseline prior constraints from the long-term causal graph to the medium-term and short-term causal graphs includes:

[0132] Extract causal pathways and their effect strengths that characterize stable seasonal patterns from long-term causal graphs;

[0133] Convert the effect intensity into a demand baseline component;

[0134] The baseline demand component is injected as a prior mean into the probability distribution parameters of the corresponding demand variables in the intermediate and short-term causal graphs;

[0135] In the cross-layer bidirectional coupling mechanism, the operation of feeding back anomalous causal strength signals from the intermediate-term causal graph and the short-term causal graph to the long-term causal graph includes:

[0136] Continuously monitor the deviation between the instantaneous strength of specific causal edges and historical baselines in medium-term and short-term causal graphs;

[0137] When the magnitude and duration of the deviation exceed a preset significance threshold, an abnormal causal strength signal is generated.

[0138] The causal strength anomaly signal and its associated contextual features are used as input events to trigger structural reassessment and parameter updates of the long-term causal graph.

[0139] In this embodiment, the cross-layer bidirectional coupling mechanism achieves bidirectional information interaction between causal graphs at different time scales by defining standardized data interfaces and event-driven logic.

[0140] The long-term causal graph transmits seasonal baseline prior constraints to the medium- and short-term causal graphs through the following process: Traversing the long-term causal graph, identifying causal paths originating from time-period variables (such as months or quarters) and pointing to core demand variables; calculating the product of the weights of each causal edge along the path to obtain the effect strength representing the overall influence of the seasonal pattern; this identification process can be automatically completed based on a graph traversal algorithm; converting the effect strength into demand baseline components using a predefined mapping relationship, which can be obtained by regression fitting by analyzing the correspondence between the effect strength of similar seasonal patterns and the actual demand baseline level in historical data; and using the obtained demand baseline components as Bayesian prior information, injecting them into the probability distribution parameters of the corresponding demand variables in the medium- and short-term causal graphs. For example, if the demand variable follows a Gaussian distribution, the baseline component is used as the mean parameter of its prior distribution, thus introducing long-term trend constraints into the probability inference of the medium- and short-term. In practice, for intermediate causal graphs, the prior mean can be directly used as the intercept term prior of the corresponding demand variable in the structural equation model; for short-term causal graphs, the mean can be used as a fixed bias prior of the output layer of the neural Granger causal network when predicting demand variables, or as an additional time-series feature input.

[0141] The intermediate-term and short-term causal graphs feed back causal strength anomaly signals to the long-term causal graph. The process is as follows: The system continuously calculates the current strength values ​​of pre-selected key causal edges in the intermediate-term and short-term causal graphs and compares them in real-time with the mean or median strength values ​​of these edges obtained statistically over a stable period (i.e., historical benchmarks, which are dynamically updated and maintained through online sliding window statistics or exponential smoothing methods), calculating the relative deviation. A preset significance threshold can be determined based on the statistical distribution of the historical deviation sequence, for example, by taking a high quantile of the absolute value of the sequence. When the deviation amplitude of a certain edge exceeds this threshold and persists for several observation periods, an anomaly is determined, and a causal strength anomaly signal is generated, containing the anomaly edge identifier, deviation value, duration stamp, and the current market snapshot. The causal strength anomaly signal and its associated contextual features are encapsulated into a specific format event message and sent to the long-term causal graph update module. The arrival of this event message constitutes the input condition for triggering structural reassessment and parameter updates in the long-term causal graph, prompting the long-term model to re-examine relevant causal hypotheses.

[0142] This embodiment, through top-down prior constraint transfer, embeds long-term patterns into the short-to-medium-term model in a parameterized form, enhancing the stability and rationality of short-to-medium-term analysis. Through bottom-up feedback of anomalous signals, the long-term model can perceive structural mutations in the market's micro-dynamics. This bidirectional coupling achieves continuous calibration and co-evolution between macro-level patterns and micro-level dynamics, making the multi-scale causal perception system an organic whole that can both inherit historical experience and adapt to environmental changes. This provides a dynamic and solid causal foundation for subsequent accurate separation of market trends and robust decision-making.

[0143] In some embodiments, potential network trends are detected early based on short-term causal graphs, and lifecycle prediction is initiated for the detected potential network trends to generate trend prediction results that include predicted demand curves and lifecycle stages, including:

[0144] Extract time-series feature indicators representing social media volume, search popularity, and information dissemination topology from short-term causal graphs;

[0145] Calculate the rate of change and acceleration of change of time-series characteristic indicators, and compare them with dynamic detection thresholds set based on historical data distribution;

[0146] When the rate of change or acceleration of change of the time sequence characteristic index exceeds the dynamic detection threshold, it is determined that a potential network storm has been detected and recorded as a storm initiation event.

[0147] In response to the trend initiation event, a hybrid lifecycle prediction model is launched;

[0148] In the hybrid life cycle prediction model, for the initial stage of the surge, the Hawkes process model is used to fit recent time series characteristic data to predict the intensity of demand stimulation in the next few hours.

[0149] In the hybrid life cycle prediction model, for the middle stage of the epidemic, a variant of the SEIR infectious disease model is used to predict the scale of infection and peak time of the potential demand group based on the simulation of node state transitions in the information propagation network.

[0150] In the hybrid life cycle prediction model, for the decline period of a hot trend, an exponential decay model is adopted, which combines the decay parameters of similar historical hot trend patterns to predict the demand decline curve.

[0151] By integrating the forecast outputs of the Hawkes process model, the SEIR infectious disease model variant, and the exponential decay model, a continuous forecast demand curve covering the entire life cycle is generated.

[0152] Based on the morphological characteristics of the continuously predicted demand curve, the initiation period, the peak period, and the decline period of the trend are divided and identified, forming a trend prediction result that includes the predicted demand curve and the life cycle stage.

[0153] In this embodiment, time-series feature indicators are extracted from the short-term causal graph. Specifically, the original time-series data corresponding to the nodes representing variables such as social media buzz and search popularity in the graph are read after cleaning and alignment. The information propagation topology is directly derived from the edge connection relationship of the short-term causal graph to form an adjacency matrix that represents the influence relationship between nodes.

[0154] The rate of change of time-series characteristic indicators is obtained by calculating the relative difference between indicator values ​​at adjacent time points, and the acceleration of change is obtained by calculating the difference between the rates of change over two consecutive time intervals. The dynamic detection threshold is set based on the distribution of historical data. For example, by statistically analyzing the numerical sequence of the rate of change and acceleration of the indicator over a past rolling time window, a specific high quantile is calculated as the threshold. This threshold is updated periodically to reflect the drift of the market baseline.

[0155] When the rate of change or acceleration of change of the time sequence characteristic indicator exceeds the above dynamic detection threshold, the system determines that a potential network storm has occurred and generates a storm initiation event record containing a timestamp, the type of triggering indicator, and the magnitude of the exceedance.

[0156] The hybrid lifecycle prediction model for the initial stage of a demand surge consists of three sub-models. The Hawkes process model is used in the early stages of the surge. This model takes anomalous pulse sequences of recent social interactions or search behaviors as input, fits the conditional strength function of the event flow through maximum likelihood estimation, predicts the intensity of event activation in the next few hours, and maps it to demand activation intensity. A variant of the SEIR infectious disease model is used in the middle stages of the surge. The contact rate matrix in the model parameters is initialized based on the information propagation topology extracted from a short-term causal graph. It numerically simulates the transfer process of four node states—susceptible, exposed, infected, and removed—on the topological network, outputting a curve showing the change in the size of the potential demand group over time and the peak arrival time. The exponential decay model is used in the decline stage of the surge. Its decay rate parameter is obtained by retrieving historical similar surge case libraries, matching the demand trajectory of the decline stage through regression fitting, and then predicting the demand decline curve.

[0157] When integrating the forecast outputs of the Hawkes process model, a variant of the SEIR infectious disease model, and the exponential decay model, the predicted values ​​of the corresponding models are used within their respective advantageous time periods. A smooth interpolation transition is performed at the time period boundaries to generate a continuous forecast demand curve covering the entire life cycle. Based on the mathematical form of this continuous forecast demand curve, life cycle stages are defined: the interval where the first derivative of the curve first turns positive and remains positive is marked as the storm surge initiation period; the interval near the extreme point where the second derivative of the curve approaches zero and the first derivative turns negative is marked as the storm surge outbreak period; and the interval where the first derivative of the curve remains negative is marked as the storm surge decline period. The final storm surge forecast result includes this continuous forecast demand curve and the corresponding time stage labels.

[0158] This embodiment achieves early detection of market trends through dynamic thresholds and employs a hybrid model for precise modeling of different stages of the trend's lifecycle. The Hawkes process captures the initial self-excitation effect, the SEIR model simulates the networked propagation in the middle stage, and exponential decay describes the decline pattern in the recession phase. Finally, smoothing integration forms a full-cycle prediction. This method transforms "trends" into concrete, phased demand curves and time markers, providing quantitative and structured input for subsequent causal signal purification and strategy adaptation, significantly improving the system's ability to perceive and predict sudden market opportunities.

[0159] In some embodiments, the credibility of edges representing causal relationships in a short-term causal graph is evaluated to generate a first causal edge credibility score, including:

[0160] For each edge representing a causal relationship in the short-term causal graph, a Granger causality test is performed to generate a Granger causality test statistic representing the stability of the time lead-lag relationship.

[0161] For social media information propagation paths associated with edges representing causal relationships, we analyze the attribute authority of nodes in the propagation path and the topological diversity of the propagation path to generate a propagation path credibility metric.

[0162] For text content associated with edges representing causal relationships, natural language processing is used to analyze the consistency of text sentiment polarity and the relevance of text topics to the core attributes of products, generating a content semantic credibility metric.

[0163] Based on Granger causality test statistics, propagation path credibility measure and content semantic credibility measure, a first causal edge credibility score for the edge representing the causal relationship is generated through weighted fusion calculation.

[0164] The strength of causal edges in the short-term causal graph is weighted and adjusted using the first causal edge credibility score, including:

[0165] The original causal strength of each edge representing a causal relationship in the short-term causal graph is multiplied by the confidence score of the first causal edge corresponding to that edge to obtain the weighted adjusted causal edge strength.

[0166] Update the weight parameters of the corresponding edges in the short-term causal graph using the weighted adjusted causal edge strength.

[0167] In this embodiment, the Granger causality test statistic is generated as follows: For a causal edge in a short-term causal graph, historical time series data corresponding to its source variable and target variable are extracted; two time series prediction models are constructed, one of which uses only the historical data of the target variable itself for prediction, and the other uses the historical data of both the source variable and the target variable for prediction; by comparing the prediction errors of the two models on the validation set, a statistical test statistic is calculated to test whether the newly added variable (historical information of the source variable) provides a significant predictive gain. This statistic is the Granger causality test statistic that characterizes the stability of the time lead-lag relationship.

[0168] When generating a credibility metric for a propagation path, the attribute authority of a node is calculated based on its publicly available attribute data, such as the number of its followers, official certification marks, and average historical content interaction rate. A pre-defined weighted scoring formula is used to derive the authority score of a single node. The node authority of a path can be the minimum or geometric mean of the authority scores of all nodes along the path. The topological diversity of a propagation path is assessed by quantifying its structural characteristics, such as the number of hops, the number of different communities to which nodes belong, or whether the path passes through key network nodes. Higher diversity generally indicates more natural and widespread propagation. The aggregated node authority value of the path is combined with the topological diversity index according to pre-defined rules, such as weighted summation or multiplication, to generate the credibility metric for the propagation path.

[0169] When generating content semantic credibility metrics, a pre-trained sentiment analysis model is used to classify the sentiment of a batch of related texts (e.g., positive, negative, neutral), and the distribution of sentiment sentiment across all texts is statistically analyzed. High consistency is characterized by a concentrated sentiment sentiment or a distribution that aligns with the expected distribution of a specific event context. For the correlation analysis between text topics and core product attributes, keywords or topic distributions are first extracted from the text and matched with a pre-constructed vocabulary of core product attributes. Relevance is quantified by calculating the semantic similarity between the text topic vector and the attribute vocabulary vector, or by statistically analyzing the frequency and contextual plausibility of attribute keywords in the text. The sentiment consistency score and topic relevance score are linearly combined according to pre-defined weights to generate the content semantic credibility metric.

[0170] The first causal edge credibility score is obtained through weighted fusion calculation. The three indicators—the Granger causality test statistic, the propagation path credibility measure, and the content semantic credibility measure—are normalized to ensure they fall within the same numerical range. A weight coefficient is assigned to each normalized indicator. These weight coefficients can be calibrated and determined by analyzing the contribution of different evidence dimensions in historical data to the final causal judgment, or pre-set by domain experts based on experience. The sum of the three weighted indicator values ​​yields the final first causal edge credibility score.

[0171] The original causal strength value of each edge representing a causal relationship in the short-term causal graph is directly multiplied by the credibility score of the first causal edge corresponding to that edge. The resulting product is the weighted adjusted causal edge strength. Using this weighted adjusted causal edge strength value, the weight parameters of the corresponding edges in the short-term causal graph are updated, thus completing the real-time correction of the model parameters.

[0172] This embodiment constructs a quantitative causal signal credibility assessment system by integrating evidence from three dimensions: time-series prediction gain test, propagation network structure analysis, and text semantic analysis. It not only examines statistical lead-lag relationships but also delves into the reliability of the information propagation path and the logical consistency of the content upon which causal relationships depend. This effectively identifies and suppresses unreliable causal associations caused by random fluctuations, false propagation, or semantic noise. The first causal edge credibility score obtained from the assessment is used to directly weight and correct the strength of causal edges in the short-term causal graph, achieving dynamic denoising and purification of the short-term causal graph model. This significantly improves the robustness and reliability of the real-time dynamic market relationships represented by the causal graph, providing a more solid and high-quality signal input for subsequent accurate prediction calibration and decision optimization.

[0173] In some embodiments, based on long-term and medium-term causality graphs, trend background calibration and operational interference filtering are performed on the storm surge forecast results to generate calibrated storm surge forecast results, including:

[0174] Extract the state of slow variable nodes representing seasonal trends and category lifecycles from the long-term causal graph, and use them as the trend background baseline;

[0175] Subtract the demand component corresponding to the trend background baseline from the continuous forecast demand curve of the wind tide forecast results to obtain the net wind tide demand curve after removing the influence of the long-term trend.

[0176] From the mid-term causal diagram, identify causal paths driven by operational activities that overlap in time with the lifecycle stages of the current trend forecast results;

[0177] Assess the strength of the impact of the causal path driven by operational activities on demand, and separate this strength from the net windfall demand curve to obtain the core windfall demand curve after removing operational disturbances;

[0178] By overlaying the trend background baseline with the influence intensity of the causal path driven by operational activities, the background and operational interference demand curves are reconstructed.

[0179] The core trend demand curve is fused with the reconstructed background and operational interference demand curves to generate a calibrated trend forecast result. The forecast demand curve in the calibrated trend forecast result has been calibrated for trend background and filtered for operational interference, and the life cycle stage identifier is inherited from the trend forecast result.

[0180] In this embodiment, the trend background baseline is extracted by analyzing the structure and parameters of a long-term causal graph. The long-term causal graph contains slow variable nodes representing time periods or product category stages. These nodes are connected to demand variables via weighted causal edges. During extraction, the state of the slow variable node corresponding to the current time point is first determined. Then, based on the weights of each edge on the causal path connecting the node to the demand variable, a preset linear or nonlinear combination function is used to calculate the baseline demand value sequence corresponding to that state. This sequence constitutes the trend background baseline. The preset combination function can be obtained by regression fitting based on the correspondence between slow variable states and demand baselines in historical data. Subtracting the trend background baseline from the continuous forecast demand curve using a time-point-by-time subtraction operation yields the net demand curve, which removes the demand portion driven by long-term, stable factors.

[0181] The system identifies causal paths for operational activities that overlap with the lifecycle stages of a trend, based on the activation time attributes of operational activity nodes in the mid-term causal graph. The mid-term causal graph records information on the changes in the activation status of different operational activities (such as promotions and advertising) as source nodes over time. The system matches the lifecycle stage time intervals of the trend prediction results with the preset or historical activation time intervals of each operational activity node in the graph, filtering out all operational activity nodes that overlap in time, and then extracting causal paths originating from these nodes.

[0182] The impact of operational activities on demand is assessed by evaluating the weights of the causal edges along the path and the intensity attributes of the operational activity nodes (such as discount level and budget level). Preferably, this is achieved by multiplying the weights of each causal edge along the path together, then multiplying by an intensity coefficient obtained from a table lookup or function mapping based on the intensity attributes of the operational activity nodes, ultimately yielding a quantified impact strength value or an intensity curve changing over time. The table lookup or mapping relationship can be established by analyzing the correlation between the weights of similar paths in historical activities and the ultimately observed demand enhancement effect. The assessed impact strength of operational activities is extracted from the net surge demand curve through numerical subtraction to obtain the core surge demand curve, reflecting the demand potential purely driven by sudden surges after removing long-term trends and known operational interferences.

[0183] The background and operational disruption demand curves are reconstructed by aligning the aforementioned trend background baseline curve with the assessed impact intensity curves of all relevant operational activities on the time axis and summing them point by point to generate a composite curve representing the normal expected demand.

[0184] When generating the calibrated wind tide forecast results, the core wind tide demand curve, which has been corrected through the above-mentioned stripping and evaluation process, is added point by point to the reconstructed background and operational disturbance demand curves to obtain a new, well-defined continuous forecast demand curve. This curve is numerically equal to the sum of the trend background, operational disturbance, and core wind tide components. Together with the life cycle stage identifier inherited from the original wind tide forecast results, it constitutes the calibrated wind tide forecast results.

[0185] This embodiment utilizes a long-term causal graph to strip away stable trend backgrounds and a medium-term causal graph to filter planned operational interferences, enabling the separation of core demand signals purely driven by sudden network surges from the original comprehensive forecast. This calibration improves the purity and accuracy of surge demand forecasts and, by reconstructing background and interference curves, achieves a clear deconstruction of the total demand forecast, clarifying the contribution of each component. The calibrated surge forecast results not only provide purer and more accurate surge demand signals but also retain the full picture of total demand by reconstructing complete curves. This allows subsequent decisions to focus on surge opportunities while also taking into account the regular business context, significantly improving the accuracy and overall rationality of replenishment strategy formulation.

[0186] In some embodiments, selecting a corresponding replenishment decision strategy based on the lifecycle stage in the calibrated storm surge forecast results includes:

[0187] Identify the lifecycle stages marked in the calibrated storm surge forecast results. The lifecycle stages include the storm surge initiation period, the storm surge peak period, and the storm surge decline period.

[0188] When the life cycle stage is the trend initiation period, the replenishment decision strategy is to choose an exploration mode strategy that takes maximizing information value as the optimization goal and adopts a high exploration rate decision mechanism.

[0189] When the life cycle stage is the peak of the trend, an aggressive mode strategy that maximizes the demand satisfaction rate and allows temporary relaxation of inventory cost constraints is chosen as the replenishment decision strategy.

[0190] When the product lifecycle stage is in the decline phase, the replenishment decision strategy is to select a clearing mode strategy that maximizes the difference between sales revenue and inventory clearance costs and links the promotion execution system.

[0191] When the calibrated trend forecast results do not identify an effective lifecycle stage, the normal mode strategy with the optimization objective of minimizing the sum of inventory holding costs and stockout loss costs is selected as the replenishment decision strategy.

[0192] In this embodiment, the lifecycle stage identified in the calibrated storm surge forecast results is determined by reading a dedicated data field in the result data structure used to record stage categories. The value of this field is generated by the aforementioned analysis of the forecast demand curve shape, explicitly marking it as the storm surge initiation period, storm surge outbreak period, or storm surge decline period. If the field is empty, or if its associated confidence score generated by the forecast model is lower than a preset threshold, it is determined that no valid lifecycle stage has been identified. This preset threshold can be determined by analyzing the relationship between the accuracy and confidence scores of historical forecast results.

[0193] When the lifecycle stage is the initial phase of a trend, the exploratory strategy aims to maximize information value. Information value here is quantified as the expected reduction in the uncertainty of future demand forecasts (typically measured by forecast variance) after replenishment. To achieve this, the exploratory strategy employs a high-exploration-rate decision-making mechanism. This means that during decision-making, actions that bring high information gains but uncertain short-term economic benefits are selected with significant probability. For example, placing a trial order far below the predicted peak demand aims to proactively collect data on the actual market reaction to revise subsequent forecasts.

[0194] When the product lifecycle stage is during a peak demand period, the aggressive strategy aims to achieve a near 100% ratio of actual available inventory to projected demand during the predicted high-demand period. This aggressive strategy allows for temporary relaxation of constraints on inventory holding costs. Specifically, this can be achieved by temporarily reducing the weighting of inventory holding costs in the total cost function during optimization calculations, or by directly adopting a higher target safety stock level predetermined through analysis of the relationship between historical stockout losses and inventory costs.

[0195] When the product lifecycle stage is in the decline phase, the clearance strategy aims to maximize the difference between sales revenue and inventory clearance costs. Sales revenue is estimated based on the predicted demand curve for the decline phase from the calibrated product trend forecast. Inventory clearance costs are calculated using a pre-set cost model that incorporates costs such as product price reduction losses, special promotional activity expenses, and additional warehousing and handling fees. The model parameters can be calibrated using historical clearance activity data. The clearance strategy works in conjunction with the promotion execution system through a pre-configured application programming interface (API), automatically converting clearance suggestions generated by the strategy decisions (such as suggested discount levels and clearance quantities) into executable instructions for the promotion system.

[0196] When no effective lifecycle stage is identified, the standard strategy aims to minimize the sum of inventory holding costs and stockout loss costs. Classic inventory control models are used for calculation, such as determining the economic order quantity or reorder point based on historical demand distribution and a pre-defined service level target. Stockout loss costs can be estimated based on potential revenue loss due to stockouts from historical sales data.

[0197] This embodiment establishes a rule system that precisely matches the lifecycle stages of a trend with replenishment decision-making strategies. For the high uncertainty of the initial phase, the strategy focuses on information gathering and market testing; for the clear high demand of the peak phase, ensuring supply is the primary task; for the inventory digestion pressure of the decline phase, maximizing clearance profits is the core objective; and for the stable period without trends, the strategy reverts to classic cost-optimal inventory management. This mapping ensures a high degree of synergy between decision-making objectives and market dynamics, transforming replenishment decisions from passive responses to proactive, intelligent choices that adapt to the core contradictions of different stages. This provides a clear and contextualized strategic framework and optimization guidance for subsequent reinforcement learning decisions, thereby significantly improving the overall effectiveness and robustness of the replenishment system in dealing with complex market changes.

[0198] In some embodiments, based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, a reinforcement learning decision module generates specific replenishment action instructions, including:

[0199] Construct a state representation for the reinforcement learning decision module. The state representation includes the current inventory level, in-transit inventory information, the predicted demand curve and life cycle stage in the calibrated windstorm forecast results, and the strength of key causal edges in the weighted and adjusted short-term causal graph.

[0200] The selected replenishment decision strategy is mapped to the policy network parameters and reward function weights of the agent in the reinforcement learning decision module, wherein the reward function includes an economic benefit term and a causal consistency reward term.

[0201] Based on the state representation, the policy network parameterized by the policy network parameters is used to generate candidate replenishment quantity actions for the target product.

[0202] Input the candidate replenishment quantity actions, status representations, weighted adjusted short-term causal graphs, long-term causal graphs, and medium-term causal graphs into the causal simulator;

[0203] By using a causal simulator, the trajectory of changes in inventory, sales and key causal variables over multiple decision-making cycles in the future is simulated after executing candidate replenishment actions.

[0204] Based on the change trajectory, the economic benefit item and the causal consistency reward item are calculated. The causal consistency reward item is generated by measuring the degree of agreement between the actual relationship between variables in the change trajectory and the causal laws revealed by the long-term causal diagram, the medium-term causal diagram, and the short-term causal diagram.

[0205] By combining the economic benefits and the causal consistency reward, we obtain an immediate reward estimate for the candidate replenishment action.

[0206] Based on the instant reward estimation and policy network parameters, the policy network parameters are updated through the policy gradient method or the actor-critic algorithm, and the final replenishment action instruction is output, which includes the replenishment quantity and the replenishment time.

[0207] In this embodiment, current inventory levels and in-transit inventory information can be directly obtained from the warehouse management system and converted into scalars or vectors. A fixed-length time series data segment is extracted from the predicted demand curve in the calibrated storm surge forecast results, and the lifecycle stages are converted into vectors using one-hot encoding. The strength of key causal edges in the weighted short-term causal graph is then extracted using a predefined set of weight values ​​corresponding to these key edges to form a strength vector. After normalization, the above data is concatenated to obtain the state representation of the reinforcement learning decision module.

[0208] Preferably, the selected replenishment decision strategy is mapped to the agent's internal parameters through a preset strategy configuration mapping table. This mapping table defines a set of corresponding configuration parameters for each replenishment decision strategy (exploration mode, aggressive mode, cleanup mode, normal mode), including the initial exploration rate parameter of the policy network and the specific weight coefficients of the economic benefit term and the causal consistency reward term in the reward function. For example, the exploration mode strategy corresponds to a high exploration rate and a higher causal consistency reward weight.

[0209] Based on this state representation, a multi-layer neural network defined by the policy network parameters is used for forward computation. With the state vector as input, after nonlinear transformation through several hidden layers, a probability distribution parameter (such as the mean and variance of a Gaussian distribution) representing the candidate replenishment quantity action for the target product is output. Then, a specific candidate replenishment quantity value is obtained by sampling or taking the expectation.

[0210] The causal simulator is a simulation environment built on a structural causal model. Based on the functional relationships or conditional probability distributions between variables defined by the input causal graphs, starting from the current state, candidate replenishment actions are used as interventions. Through numerical iteration, the sequence of state changes of inventory, sales volume and other key market variables (such as social media buzz) in multiple future time steps after the replenishment action is executed is deduced, i.e., the change trajectory.

[0211] Based on the simulated trajectory, the economic benefit is calculated by adding the cumulative sales revenue along the trajectory and subtracting the holding costs, potential stockout losses, and replenishment costs based on inventory levels. The causal consistency reward is generated by quantifying the match between the statistical relationships (such as covariance and lead-lag correlation coefficients) between the variables in the trajectory and the causal laws encoded by the input causal graph. For example, it checks whether the outcome variable also shows a corresponding trend change when the causal variable increases, and whether the magnitude of the change roughly matches the weights of the edges in the causal graph. The higher the match, the greater the reward value.

[0212] The economic benefit and causal consistency reward terms are combined, and each is multiplied by a weighting coefficient determined by the selected strategy and then summed to obtain an immediate reward estimate for the candidate replenishment action.

[0213] Based on this immediate reward estimation, the parameters of the policy network are updated using either the policy gradient method or the actor-critic algorithm. The policy gradient method directly calculates the gradient with respect to the policy parameters using the reward signal and updates along the gradient direction. The actor-critic algorithm, on the other hand, simultaneously maintains a critic network to evaluate the value of states or actions, using value estimation to guide the updates of the policy network (actors) to reduce variance and improve learning stability. After one round of policy optimization, either online or offline, the action output by the current policy network processing the initial state is determined as the final replenishment action instruction. This instruction explicitly includes the suggested replenishment quantity and the suggested time point for executing the replenishment operation.

[0214] This embodiment integrates multi-scale causal graphs into state representation and a causal simulator, enabling the agent to perceive and utilize the deep causal structure of the market. By designing a reward function that includes a causal consistency reward term, the agent is guided to learn replenishment strategies that are not only cost-effective but also conform to market causal laws. Through trajectory extrapolation based on the causal simulator, the long-term impact can be assessed before taking actual actions, achieving deliberate decision-making. This embodiment combines the exploratory optimization capabilities of reinforcement learning with the interpretability and counterfactual reasoning capabilities of causal models, making the generated replenishment action instructions not only adaptable to different market stages (through policy mapping) but also possessing a solid causal logical foundation. This significantly improves the scientific rigor, robustness, and long-term profit potential of replenishment decisions in complex and dynamic market environments.

[0215] In some embodiments, parameters of the long-term causal graph, medium-term causal graph, and short-term causal graph are updated online using actual sales and inventory feedback data, and the policy parameters of the reinforcement learning decision module are optimized based on feedback, including:

[0216] The actual sales and inventory feedback data are compared with the predicted demand curve in the calibrated trend forecast results to generate a prediction error sequence;

[0217] Based on the prediction error sequence, online gradient descent is performed to update the weight parameters of relevant causal edges in the short-term causal graph to reduce the prediction error sequence.

[0218] Separate the systematic deviation components in the prediction error sequence that are related to long-term trends and operational activities;

[0219] Using the systematic bias component, Bayesian updates are performed on the node parameters representing the seasonal baseline in the long-term causal graph and the causal edge strength parameters representing the impact of operational activities in the medium-term causal graph, respectively.

[0220] The actual sales and inventory feedback data, the replenishment action instructions executed, and the updated long-term, medium-term, and short-term causal graphs together constitute the experience sample for the reinforcement learning decision-making module.

[0221] Store the experience samples into the experience replay buffer;

[0222] Sample batches of experience samples from the experience replay buffer and calculate the advantage function estimate for each state-action pair in the batch of experience samples.

[0223] Based on advantage function estimation, the policy network parameters and value network parameters of the agent in the reinforcement learning decision module are iteratively updated using the proximal policy optimization algorithm or the soft actor-critic algorithm to complete feedback optimization.

[0224] In this embodiment, the prediction error sequence is obtained by subtracting the actual observed sales and inventory data from the predicted demand values ​​at the corresponding time points in the calibrated storm surge prediction results point by point.

[0225] Online gradient descent updates are performed on the weight parameters of relevant causal edges in the short-term causal graph, with the prediction error sequence as a supervision signal. The sum of squares of the prediction error or other forms of loss function are used to calculate the gradient of the loss with respect to the weight parameters of the causal edges through the backpropagation algorithm. Then, these weights are adjusted in the opposite direction of the gradient according to the preset learning rate, so that the prediction output by the model is closer to the actual observation.

[0226] To separate the systematic deviation components in the forecast error series that are related to long-term trends and operational activities, time series decomposition techniques can be used. For example, the error series can be decomposed into trend, seasonal, cyclical, and residual components, among which the trend component and the cyclical component that is synchronized with the known operational activity cycle are identified as systematic deviations.

[0227] Using the isolated systematic bias components, Bayesian updates are performed on the node parameters representing the seasonal baseline in the long-term causal graph. The trend or seasonal pattern in the systematic bias is used as new evidence to update the parameters of the probability distribution (such as the mean of a Gaussian distribution) followed by the node. Similarly, Bayesian updates are performed on the causal edge strength parameters representing the impact of operational activities in the medium-term causal graph. The components of the systematic bias that match the time of a specific operational activity are used as evidence to correct the posterior distribution of the causal edge strength parameter.

[0228] The actual sales and inventory feedback data, the executed replenishment instructions, and the updated long-term, medium-term, and short-term cause-effect graphs are collectively encapsulated into a structured experience sample. This sample fully records the state, actions, rewards (which can be derived from feedback data), and new state of a single decision loop. The encapsulated experience sample is stored in an experience replay buffer with a first-in-first-out (FIFO) or priority sorting mechanism. This buffer is used to store historical experience for subsequent batch learning.

[0229] A batch of experience samples is randomly sampled from the experience replay buffer. Using the current value network or through a generalized advantage estimation algorithm, the advantage function estimate corresponding to the state-action pair in each sample is calculated. This value quantifies the superiority or inferiority of a particular action relative to the average level.

[0230] Based on the calculated advantage function estimate, either the proximal policy optimization algorithm or the soft actor-critic algorithm is used to iteratively update the parameters of the policy network and value network in the reinforcement learning decision module. Proximal policy optimization stabilizes policy updates by optimizing an alternative objective function with pruning; the soft actor-critic algorithm maximizes policy entropy while maximizing expected reward, thus balancing exploration and exploitation. Both algorithms utilize advantage function estimation to guide the adjustment of network parameters, thereby improving policy performance.

[0231] This embodiment dynamically corrects the parameters of the multi-scale causal graph using gradient descent and Bayesian methods, enabling the market perception model to track changes. Experience replay and strategy optimization algorithms continuously improve the policy and value networks of the decision-making agent. The synergistic optimization of the perception model and decision-making strategies ensures that the system can simultaneously improve its understanding of market patterns and the effectiveness of its actions based on real-time feedback, thereby achieving long-term, stable adaptive performance improvement in the complex and ever-changing cross-border e-commerce environment.

[0232] Please see Figure 3In a second aspect, this embodiment also provides a cross-border e-commerce adaptive replenishment system 1 based on multi-scale causal perception, applicable to the method described in the first aspect. The system includes a data fusion module 11, a multi-scale causal graph construction and update module 12, a trend detection and prediction module 13, a causal signal credibility assessment module 14, a trend prediction calibration module 15, a strategy selection module 16, a reinforcement learning decision-making module 17, a decision execution and feedback module 18, and an online learning and optimization module 19. The data fusion module 11 is used to receive multi-source heterogeneous real-time data streams and perform time alignment and... Feature fusion generates a unified multi-dimensional temporal feature vector; the multi-scale causal graph construction and update module 12 is used to construct long-term, medium-term, and short-term causal graphs in parallel based on the multi-dimensional temporal feature vector, and dynamically update the long-term, medium-term, and short-term causal graphs, which interact through cross-layer coupling; the trend detection and prediction module 13 is used to perform early detection of potential network trends based on the short-term causal graph, and initiate lifecycle prediction for the detected potential network trends, generating trend prediction results that include the predicted demand curve and lifecycle stage; causal... The signal credibility assessment module 14 is used to assess the credibility of edges representing causal relationships in the short-term causal graph, generate a first causal edge credibility score, and use the first causal edge credibility score to weight and adjust the strength of causal edges in the short-term causal graph; the storm surge prediction calibration module 15 is used to perform trend background calibration and operational interference filtering on the storm surge prediction results based on the long-term and medium-term causal graphs, and generate calibrated storm surge prediction results; the strategy selection module 16 is used to select the corresponding replenishment decision strategy according to the life cycle stage in the calibrated storm surge prediction results. The replenishment decision strategy includes normal mode strategy, exploration... The reinforcement learning decision module 17 is used to generate specific replenishment action instructions based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend forecast results. The decision execution and feedback module 18 is used to execute the replenishment action instructions and collect actual sales and inventory feedback data. The online learning and optimization module 19 is used to update the parameters of the long-term, medium-term, and short-term causal graphs online using actual sales and inventory feedback data, and to provide feedback optimization to the strategy parameters of the reinforcement learning decision module 17.

[0233] Through the coordinated operation of the aforementioned modules, this system transforms multi-source heterogeneous data streams into multi-scale market cognition with causal interpretability. Based on this cognition, it drives reinforcement learning agents to generate adaptive replenishment decisions. After executing a decision, the system uses real market feedback to simultaneously update its cognitive model (causal graph) and decision model (policy network), forming a complete closed loop from perception, decision-making, execution to learning. This not only enables a rapid and accurate response to sudden network fluctuations but also allows for continuous evolution in long-term operation, constantly improving the intelligence level and overall robustness of replenishment decisions in the complex and ever-changing cross-border e-commerce environment.

[0234] By adopting the above technical solutions, this invention differs from existing technologies and possesses the following beneficial effects: By constructing and coupling long-term, medium-term, and short-term causal graphs in parallel across layers, it achieves collaborative causal perception and decoupling of multi-scale market dynamics, effectively separating sudden network surge signals from long-term trend backgrounds and medium-term operational interference; by conducting multi-dimensional credibility assessment and weighted adjustment of causal signals in the short-term causal graph, it enhances the robustness of real-time market dynamic perception; furthermore, it adaptively selects replenishment decision strategies based on the lifecycle stage of the calibrated surge prediction results, and uses a reinforcement learning module incorporating causal constraints to generate replenishment action instructions, enabling decisions to accurately match the core contradictions of different market stages; finally, it collaboratively optimizes the causal graph model and decision strategies online using actual feedback data, forming a complete closed loop of perception, decision-making, execution, and learning, significantly improving the early detection and accurate prediction capabilities of sudden demands in complex and non-stationary market environments, as well as the adaptability and long-term robustness of replenishment decisions.

[0235] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0236] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0237] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A cross-border e-commerce adaptive replenishment method based on multi-scale causal perception, characterized in that, include: Receive multi-source heterogeneous real-time data streams; Time alignment and feature fusion are performed on multi-source heterogeneous real-time data streams to generate a unified multi-dimensional time-series feature vector; Based on multidimensional temporal feature vectors, long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated. The long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs interact through cross-layer coupling. Early detection of potential network trends is performed based on short-term causal graphs, and lifecycle prediction is initiated for the detected potential network trends to generate trend prediction results that include predicted demand curves and lifecycle stages. The credibility of edges representing causal relationships in short-term causal graphs is evaluated to generate a first causal edge credibility score, and the strength of causal edges in short-term causal graphs is weighted and adjusted using the first causal edge credibility score. Based on long-term and medium-term causal graphs, trend background calibration and operational interference filtering are performed on the storm surge forecast results to generate calibrated storm surge forecast results. The corresponding replenishment decision strategy is selected based on the life cycle stage in the calibrated wind tide forecast results. The replenishment decision strategy includes one of the following: normal mode strategy, exploratory mode strategy, aggressive mode strategy, and cleanup mode strategy. Based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, specific replenishment action instructions are generated through the reinforcement learning decision module. Execute replenishment orders and collect actual sales and inventory feedback data; Using actual sales and inventory feedback data, the parameters of the long-term, medium-term, and short-term causal graphs are updated online, and the strategy parameters of the reinforcement learning decision-making module are optimized based on feedback. Specifically, based on multi-dimensional temporal feature vectors, long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs are constructed in parallel and dynamically updated, including: Data aggregation at different time scales is performed on multidimensional time series feature vectors to generate long-term aggregated feature vectors, medium-term aggregated feature vectors, and short-term aggregated feature vectors. Based on long-term aggregated feature vectors, the causal structure between monthly or quarterly scale variables is discovered through the FCI algorithm, and a long-term causal graph is constructed. Based on the mid-term aggregated feature vector and combined with the pre-set prior knowledge of operational activities, a mid-term causal graph representing the causal influence between weekly operational variables is constructed using the LiNGAM algorithm or a constraint-based causal discovery algorithm. Based on short-term aggregated feature vectors, the instantaneous causal relationships between variables in high-dimensional streaming data at the minute or hour scale are discovered through the neural Granger causality algorithm, and a short-term causal graph is constructed.

2. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception as described in claim 1, characterized in that, Long-term causal graphs, intermediate-term causal graphs, and short-term causal graphs interact through cross-layer coupling, including: Establish a cross-layer bidirectional coupling mechanism between long-term causal graphs, medium-term causal graphs, and short-term causal graphs; In the aforementioned cross-layer bidirectional coupling mechanism, the operation of transferring seasonal baseline prior constraints from the long-term causal graph to the medium-term and short-term causal graphs includes: Extract causal pathways and their effect strengths that characterize stable seasonal patterns from long-term causal graphs; Convert the effect intensity into a demand baseline component; The baseline demand component is injected as a prior mean into the probability distribution parameters of the corresponding demand variables in the intermediate-term causal graph and the short-term causal graph; In the aforementioned cross-layer bidirectional coupling mechanism, the operation of feeding back anomalous causal intensity signals from the intermediate-term causal graph and the short-term causal graph to the long-term causal graph includes: Continuously monitor the deviation between the instantaneous strength of specific causal edges and historical baselines in medium-term and short-term causal graphs; When the magnitude and duration of the deviation exceed a preset significance threshold, a causal strength anomalous signal is generated; The causal strength anomaly signal and its associated contextual features are used as input events to trigger structural reassessment and parameter updates of the long-term causal graph.

3. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception according to claim 1, characterized in that, Early detection of potential network trends is performed based on short-term causal graphs. Lifecycle prediction is then initiated for the detected potential network trends, generating trend prediction results that include predicted demand curves and lifecycle stages. Extract time-series feature indicators representing social media volume, search popularity, and information dissemination topology from short-term causal graphs; Calculate the rate of change and acceleration of change of the time-series characteristic indicators, and compare them with the dynamic detection threshold set based on the historical data distribution; When the rate of change or acceleration of change of the time-series characteristic indicator exceeds the dynamic detection threshold, it is determined that a potential network storm has been detected and recorded as a storm initiation event. In response to the aforementioned trend initiation event, a hybrid lifecycle prediction model is activated; In the hybrid life cycle prediction model, for the initial stage of the demand surge, the Hawkes process model is used to fit recent time-series characteristic data to predict the demand surge intensity in the next few hours. In the hybrid life cycle prediction model, for the middle stage of the epidemic, a variant of the SEIR infectious disease model is used to predict the scale of infection and peak time of the potential demand group based on the simulation of node state transitions in the information propagation network. In the hybrid life cycle prediction model, for the wind tide decline period, an exponential decay model is adopted, combined with the decay parameters of similar historical wind tide patterns, to predict the demand decline curve. By integrating the prediction outputs of the Hawkes process model, the SEIR infectious disease model variant, and the exponential decay model, a continuous forecast demand curve covering the entire life cycle is generated. Based on the morphological characteristics of the continuous forecast demand curve, the storm initiation period, storm outbreak period, and storm decline period are divided and identified, forming a storm forecast result that includes the forecast demand curve and the life cycle stage.

4. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception as described in claim 1, characterized in that, The credibility of edges representing causal relationships in a short-term causal graph is evaluated to generate a first causal edge credibility score, including: For each edge representing a causal relationship in the short-term causal graph, a Granger causality test is performed to generate a Granger causality test statistic representing the stability of the time lead-lag relationship. For social media information propagation paths associated with the edges representing causal relationships, the attribute authority of nodes in the propagation path and the topological diversity of the propagation path are analyzed to generate a propagation path credibility metric. For text content associated with the edges representing causal relationships, natural language processing is used to analyze the consistency of text sentiment polarity and the relevance of text topics to the core attributes of products, generating a content semantic credibility metric. Based on the Granger causality test statistic, the propagation path credibility metric, and the content semantic credibility metric, a first causal edge credibility score for the edge representing the causal relationship is generated through weighted fusion calculation. The strength of causal edges in the short-term causal graph is weighted and adjusted using the first causal edge credibility score, including: The original causal strength of each edge representing a causal relationship in the short-term causal graph is multiplied by the confidence score of the first causal edge corresponding to that edge to obtain the weighted adjusted causal edge strength. The weight parameters of the corresponding edges in the short-term causal graph are updated using the weighted adjusted causal edge strength.

5. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception according to claim 1, characterized in that, Based on long-term and medium-term causal graphs, trend background calibration and operational interference filtering are applied to the storm surge forecast results, generating calibrated storm surge forecast results, including: Extract the state of slow variable nodes representing seasonal trends and category lifecycles from the long-term causal graph, and use them as the trend background baseline; Subtract the demand component corresponding to the trend background baseline from the continuous forecast demand curve of the wind tide forecast results to obtain the net wind tide demand curve after removing the influence of the long-term trend. From the mid-term causal diagram, identify causal paths driven by operational activities that overlap in time with the lifecycle stages of the current trend forecast results; Assess the impact strength of the causal path driven by operational activities on demand, and separate this impact strength from the net demand curve to obtain the core demand curve after removing operational interference; The background trend baseline is superimposed with the influence intensity of the causal path driven by operational activities to reconstruct the background and operational interference demand curves. The core wind tide demand curve is fused with the reconstructed background and operational interference demand curves to generate a calibrated wind tide forecast result. The forecast demand curve in the calibrated wind tide forecast result has been calibrated for trend background and filtered for operational interference, and the life cycle stage identifier is inherited from the wind tide forecast result.

6. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception according to claim 1, characterized in that, Based on the lifecycle stage in the calibrated storm surge forecast results, select the corresponding replenishment decision strategy, including: Identify the life cycle stages marked in the calibrated wind tide forecast results, which include the wind tide initiation period, the wind tide outbreak period, and the wind tide decline period. When the life cycle stage is the trend initiation period, the replenishment decision strategy is to select an exploration mode strategy that takes maximizing information value as the optimization goal and adopts a high exploration rate decision mechanism. When the life cycle stage is the peak period of a trend, an aggressive mode strategy that maximizes the demand satisfaction rate and allows temporary relaxation of inventory cost constraints is selected as the replenishment decision strategy. When the life cycle stage is the decline period, the replenishment decision strategy is to select a clearing mode strategy that maximizes the difference between sales revenue and inventory clearance cost and links the promotion execution system. When the calibrated trend forecast results do not identify an effective lifecycle stage, the normal mode strategy with the optimization objective of minimizing the sum of inventory holding costs and stockout loss costs is selected as the replenishment decision strategy.

7. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception according to claim 1, characterized in that, Based on the selected replenishment decision strategy, the weighted adjusted short-term causal graph, and the calibrated trend prediction results, the reinforcement learning decision module generates specific replenishment action instructions, including: A state representation for a reinforcement learning decision-making module is constructed, which includes the current inventory level, in-transit inventory information, the predicted demand curve and life cycle stage in the calibrated storm forecast results, and the strength of key causal edges in the weighted and adjusted short-term causal graph. The selected replenishment decision strategy is mapped to the policy network parameters and reward function weights of the agent in the reinforcement learning decision module, wherein the reward function includes an economic benefit term and a causal consistency reward term. Based on the state representation, a candidate replenishment quantity action for the target product is generated using the policy network parameterized by the policy network parameters. The candidate replenishment quantity action, the state representation, and the weighted adjusted short-term causal graph, long-term causal graph, and medium-term causal graph are all input into the causal simulator. The causal simulator is used to simulate the changes in inventory, sales volume, and key causal variables over multiple future decision-making cycles after the candidate replenishment action is executed. Based on the change trajectory, the economic benefit item and the causal consistency reward item are calculated, wherein the causal consistency reward item is generated by measuring the degree of agreement between the actual relationship between variables in the change trajectory and the causal laws revealed by the long-term causal diagram, the medium-term causal diagram, and the short-term causal diagram. By combining the economic benefit item and the causal consistency reward item, an immediate reward estimate for the candidate replenishment quantity action is obtained; Based on the instant reward estimate and the policy network parameters, the policy network parameters are updated using the policy gradient method or the actor-critic algorithm, and the final replenishment action instruction is output, which includes the replenishment quantity and the replenishment time.

8. The adaptive replenishment method for cross-border e-commerce based on multi-scale causal perception according to claim 1, characterized in that, Using actual sales and inventory feedback data, the parameters of the long-term, medium-term, and short-term causal graphs are updated online, and the policy parameters of the reinforcement learning decision-making module are optimized based on feedback, including: The actual sales and inventory feedback data are compared with the predicted demand curve in the calibrated trend forecast results to generate a prediction error sequence; Based on the prediction error sequence, online gradient descent is performed to update the weight parameters of the relevant causal edges in the short-term causal graph to reduce the prediction error sequence. The systematic deviation components related to long-term trends and operational activities in the prediction error sequence are separated; Using the aforementioned systematic bias components, Bayesian updates are performed on the node parameters representing the seasonal baseline in the long-term causal graph and the causal edge strength parameters representing the impact of operational activities in the medium-term causal graph, respectively. The actual sales and inventory feedback data, the replenishment action instructions executed, and the updated long-term, medium-term, and short-term causal graphs together constitute the experience sample for the reinforcement learning decision-making module. The experience samples are stored in the experience replay buffer; Sample batches of experience samples from the experience replay buffer and calculate the advantage function estimate for each state-action pair in the batches of experience samples; Based on the aforementioned advantage function estimation, the policy network parameters and value network parameters of the agent in the reinforcement learning decision module are iteratively updated using a proximal policy optimization algorithm or a soft actor-critic algorithm to complete feedback optimization.

9. A cross-border e-commerce adaptive replenishment system based on multi-scale causal perception, characterized in that, The system applicable to the method of any one of claims 1 to 8, the system comprising: The data fusion module is used to receive multi-source heterogeneous real-time data streams, and to perform time alignment and feature fusion on the multi-source heterogeneous real-time data streams to generate a unified multi-dimensional time-series feature vector. The multi-scale causal graph construction and update module is used to construct long-term, medium-term, and short-term causal graphs in parallel based on multi-dimensional temporal feature vectors, and dynamically update the long-term, medium-term, and short-term causal graphs. The long-term, medium-term, and short-term causal graphs interact through cross-layer coupling. The trend detection and prediction module is used to detect potential network trends in the early stage based on short-term causal graphs, and to initiate life cycle prediction for the detected potential network trends, generating trend prediction results that include the predicted demand curve and life cycle stage. The causal signal credibility assessment module is used to assess the credibility of edges representing causal relationships in short-term causal graphs, generate a first causal edge credibility score, and use the first causal edge credibility score to weight and adjust the strength of causal edges in short-term causal graphs. The storm surge forecast calibration module is used to perform trend background calibration and operational interference filtering on the storm surge forecast results based on long-term and medium-term causal graphs, and generate calibrated storm surge forecast results. The strategy selection module is used to select the corresponding replenishment decision strategy based on the life cycle stage in the calibrated wind tide forecast results. The replenishment decision strategy includes one of the following: normal mode strategy, exploratory mode strategy, aggressive mode strategy, and cleanup mode strategy. The reinforcement learning decision-making module is used to generate specific replenishment action instructions based on the selected replenishment decision-making strategy, the weighted adjusted short-term causal graph, and the calibrated trend forecast results. The decision execution and feedback module is used to execute replenishment instructions and collect actual sales and inventory feedback data. The online learning and optimization module is used to update the parameters of the long-term, medium-term, and short-term causal graphs online using actual sales and inventory feedback data, and to provide feedback optimization for the strategy parameters of the reinforcement learning decision-making module.