Order quantity prediction method and device, electronic equipment and storage medium

By employing techniques such as multi-scale causal attention networks and variational dynamic hybrid modeling, and combining multi-source data for order volume prediction, this approach solves the prediction accuracy problems of traditional methods in capturing multi-dimensional factors and nonlinear relationships. It achieves high-precision and highly interpretable order volume prediction, supporting enterprise decision-making in complex market environments.

CN120912253APending Publication Date: 2025-11-07CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511021959.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional order volume forecasting methods cannot fully capture multidimensional factors, resulting in a significant decrease in forecast accuracy and robustness when faced with data noise, sudden events, or nonlinear relationships, and they lack high-precision and highly interpretable forecasting solutions.

Method used

By acquiring historical order data, advertising data, social media data, and real-time exogenous variables, and after data alignment and cleaning, a multimodal feature matrix is ​​extracted. Wavelet decomposition and causal attention calculation are performed using a multi-scale causal attention network. Combined with variational dynamic hybrid modeling, dynamic adaptive embedding mechanism, and adversarial reinforcement learning, adaptive causal intervention analysis is conducted, ultimately obtaining the target order volume prediction sequence and causal effect estimation.

Benefits of technology

It improves the comprehensiveness and accuracy of order volume forecasting, providing strong support for decisions such as marketing strategies and warehouse allocation, and helping companies maintain a competitive advantage in a complex and ever-changing market environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912253A_ABST
    Figure CN120912253A_ABST
Patent Text Reader

Abstract

The invention discloses an order quantity prediction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining historical order data, advertisement promotion data, social media data and real-time exogenous variables, and obtaining a multi-modal feature matrix through data alignment, data cleaning and feature extraction; performing multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix to obtain an initial order quantity prediction sequence and an attention weight matrix; variational dynamic hybrid modeling and variational reasoning optimization are carried out on the initial order quantity prediction sequence to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters; carrying out adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters to obtain refined feature embedding; and performing adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence and refined feature embedding to obtain a target order quantity prediction sequence. The method improves the comprehensiveness and accuracy of product order quantity prediction, and can be widely applied to the technical field of information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, and in particular to an order quantity prediction method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of digital marketing, the market environment faced by enterprises is becoming increasingly complex and variable. The slowdown of economic growth has led to increasingly fierce competition in the communication market, and the prediction of product order quantity has become a key link for optimizing the supply chain, formulating marketing strategies and improving competitiveness.

[0003] Traditional prediction methods, such as ARIMA models based on single time series or simple machine learning regression, often fail to fully capture the multi-dimensional factors affecting order quantity, such as the dynamic effect of advertising, the public opinion fluctuation of social media, seasonal trends and the strategy adjustment of competitors. In addition, when facing data noise, sudden events or nonlinear relationships, the prediction accuracy and robustness of these methods decrease significantly. At the same time, the demand for prediction systems has shifted from simple numerical output to more business value insights, such as uncertainty quantification, causal effect analysis and real-time adaptability. Therefore, it is particularly urgent to design a prediction scheme that integrates multi-source data, has high accuracy and high interpretability. SUMMARY

[0004] The present application aims to at least partially solve one of the problems in the prior art.

[0005] To this end, one object of the present application is to provide an order quantity prediction method that improves the comprehensiveness and accuracy of product order quantity prediction.

[0006] Another object of the present application is to provide an order quantity prediction device.

[0007] In order to achieve the above technical purpose, the technical solution adopted by the embodiments of the present application comprises:

[0008] On the one hand, the embodiments of the present application provide an order quantity prediction method, comprising the following steps:

[0009] Obtain historical order data, advertising promotion data, social media data and real-time exogenous variables, and obtain multi-source time series data through data alignment and data cleaning, and obtain a multi-modal feature matrix by feature extraction on the multi-source time series data;

[0010] Perform multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network to obtain an initial order quantity prediction sequence and an attention weight matrix;

[0011] The initial order quantity prediction sequence is subjected to variational dynamic hybrid modeling and variational inference optimization to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters;

[0012] The hidden variable distribution parameters are subjected to adversarial reinforcement learning and feature reconstruction based on a dynamic adaptive embedding mechanism to obtain refined feature embeddings;

[0013] The attention weight matrix, the optimized order quantity prediction sequence, and the refined feature embeddings are subjected to adaptive causal intervention analysis to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable.

[0014] Further, in an embodiment of the present application, the multi-source time series data is obtained through data alignment and data cleaning, and multi-modal feature matrices are extracted from the multi-source time series data, which specifically includes:

[0015] The historical order data, the advertising promotion data, the social media data, and the real-time exogenous variables are subjected to dynamic time warping for time alignment, and missing value filling and outlier removal to obtain the multi-source time series data;

[0016] The multi-source time series data is subjected to kernel principal component analysis to extract the multi-modal feature matrices.

[0017] Further, in an embodiment of the present application, the multi-modal feature matrices are subjected to multi-scale wavelet decomposition and causal attention calculation through a multi-scale causal attention network to obtain an initial order quantity prediction sequence and an attention weight matrix, which specifically includes:

[0018] The multi-modal feature matrices are decomposed into short-period components, medium-period components, and long-period components through discrete wavelet transformation;

[0019] The short-period components, the medium-period components, and the long-period components are subjected to hidden state calculation and information transmission through a gated recurrent unit to obtain short-period prediction sequences, medium-period prediction sequences, and long-period prediction sequences;

[0020] The attention weights of the short-period components, the medium-period components, and the long-period components are calculated based on a causal attention mechanism to obtain the attention weight matrix;

[0021] The short-period prediction sequences, the medium-period prediction sequences, and the long-period prediction sequences are subjected to dynamic weighted fusion according to the attention weight matrix to obtain the initial order quantity prediction sequence.

[0022] Further, in one embodiment of the present application, the initial order quantity prediction sequence is variational dynamic mixture modeled and variational inference optimized to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters, which specifically comprises:

[0023] An initial variational distribution is constructed according to the initial order quantity sequence, and the initial variational distribution is input into an uncertainty network to estimate an order quantity mean and an order quantity variance;

[0024] The initial order quantity sequence, the order quantity mean, and the order quantity variance are modeled according to a Gaussian mixture model to obtain a dynamic Gaussian mixture distribution;

[0025] The dynamic Gaussian mixture distribution is obtained by minimizing the KL divergence loss through variational inference based on the approximate posterior distribution;

[0026] The optimized order quantity prediction sequence, the corresponding optimized order quantity mean, and the optimized order quantity variance are determined according to the optimized dynamic Gaussian mixture distribution, and the hidden variable distribution parameters are generated according to the optimized order quantity mean and the optimized order quantity variance.

[0027] Further, in one embodiment of the present application, the hidden variable distribution parameters are adversarially enhanced and reconstructed based on a dynamic adaptive embedding mechanism to obtain refined feature embeddings, which specifically comprises:

[0028] A dynamic pseudo-distribution parameter is generated by a time-conditioned generator according to the historical context of the hidden variable distribution parameters, and a distribution consistency loss of the dynamic pseudo-distribution parameter and the hidden variable distribution parameters is judged by a discriminator;

[0029] A dynamic pseudo-order quantity prediction sequence is determined according to the dynamic pseudo-distribution parameter, and a prediction consistency loss of the dynamic pseudo-order quantity prediction sequence and the optimized order quantity prediction sequence is determined;

[0030] A joint loss of adversarial training is determined according to the distribution consistency loss and the prediction consistency loss, and the joint loss is optimized through deep reinforcement learning to obtain a trained time-conditioned generator;

[0031] A target distribution parameter is generated by the trained time-conditioned generator, and the target distribution parameter is reconstructed to obtain the refined feature embeddings.

[0032] Further, in one embodiment of the present application, the attention weight matrix, the optimized order quantity prediction sequence, and the refined feature embeddings are adaptively and causally intervened to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable, which specifically comprises:

[0033] construct a dynamic causal graph with time series features of each modality of the multi-modal feature matrix as nodes, and update edge weights between nodes through Granger causality test and the attention weight matrix;

[0034] calculate the causal effect estimation value of each intervention variable through counterfactual reasoning according to the refined feature embedding and the dynamic causal graph;

[0035] update the optimized order quantity prediction sequence according to the causal effect estimation value to obtain the target order quantity prediction sequence.

[0036] Further, in an embodiment of the present application, the order quantity prediction method further comprises the following steps:

[0037] obtain an order quantity real sequence of a target period, and determine a real order quantity mean and a real order quantity variance according to the order quantity real sequence;

[0038] calculate an order quantity prediction loss according to the order quantity real sequence and the target order quantity prediction sequence;

[0039] calculate an evidence lower bound loss according to the order quantity real sequence and the optimized dynamic Gaussian mixture distribution;

[0040] calculate a real distribution loss according to the real order quantity mean, the real order quantity variance, and the hidden variable distribution parameter;

[0041] calculate a causal prediction loss according to a change rate of the order quantity real sequence at each time point and a change rate of the target order quantity prediction sequence at each time point;

[0042] determine an overall loss value according to the order quantity prediction loss, the evidence lower bound loss, the real distribution loss, and the causal prediction loss;

[0043] update parameters of the multi-scale causal attention network, the variational dynamic mixture modeling, the adversarial reinforcement learning, and the adaptive causal intervention analysis according to the overall loss value.

[0044] In another aspect, an embodiment of the present application provides an order quantity prediction device, comprising:

[0045] a data acquisition and processing module, configured to obtain historical order data, advertising promotion data, social media data, and real-time exogenous variables, and obtain multi-source time series data through data alignment and data cleaning, and obtain a multi-modal feature matrix through feature extraction on the multi-source time series data;

[0046] a multi-scale causal attention module, configured to perform multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network, to obtain an initial order quantity prediction sequence and an attention weight matrix;

[0047] a variational dynamic hybrid modeling module, configured to perform variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence, to obtain an optimized order quantity prediction sequence and a latent variable distribution parameter;

[0048] a dynamic adaptive embedding learning module, configured to perform adversarial reinforcement learning and feature reconstruction on the latent variable distribution parameter based on a dynamic adaptive embedding mechanism, to obtain a refined feature embedding;

[0049] an adaptive causal intervention analysis module, configured to perform adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence and the refined feature embedding, to obtain a target order quantity prediction sequence and a causal effect estimation value of each intervention variable.

[0050] In another aspect, an embodiment of the present application provides an electronic device, which comprises a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and the program is executed by the processor to realize the order quantity prediction method as described above.

[0051] In another aspect, an embodiment of the present application further provides a storage medium, which is a computer readable storage medium, configured to be computer readable, and the storage medium stores one or more programs, and the one or more programs are executable by one or more processors to realize the order quantity prediction method as described above.

[0052] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present application:

[0053] The embodiment of the application obtains historical order data, advertisement promotion data, social media data and real-time exogenous variables, obtains multi-source time series data through data alignment and data cleaning, obtains a multi-modal feature matrix through feature extraction on the multi-source time series data, performs multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network, obtains an initial order quantity prediction sequence and an attention weight matrix, performs variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence, obtains an optimized order quantity prediction sequence and hidden variable distribution parameters, performs adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters based on a dynamic adaptive embedding mechanism, obtains refined feature embedding, and performs adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence and the refined feature embedding, to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable. The embodiment of the application fuses multi-modal data such as historical orders, advertisement promotion, social media dynamics and exogenous variables, realizes multi-level feature extraction and order quantity prediction through multi-scale causal attention network fusion wavelet transformation and causal masking, obtains an optimized order quantity prediction sequence by quantifying the uncertainty of prediction through variational dynamic hybrid modeling and variational inference, estimates the intervention effect of each intervention variable on the order quantity through adversarial reinforcement learning and adaptive causal intervention analysis, and finally obtains the target order quantity prediction sequence, thereby improving the comprehensiveness and accuracy of product order quantity prediction, providing strong support for marketing strategies, warehouse allocation and other decisions, and effectively helping enterprises maintain a competitive advantage in a complex and changing market environment. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following introduces the drawings needed to be used in the embodiments of the present application. It should be understood that the drawings introduced below are only for facilitating clear description of some embodiments in the technical solutions of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the premise that the drawings are not provided.

[0055] Figure 1 A step flow chart of the order quantity prediction method provided by the embodiment of the present application;

[0056] Figure 2 A step flow chart of step S101 provided by the embodiment of the present application;

[0057] Figure 3 A step flow chart of step S102 provided by the embodiment of the present application;

[0058] Figure 4 A step flow chart of step S103 provided by the embodiment of the present application

[0059] Figure 5A step flow chart of step S104 provided for the embodiment of the present application;

[0060] Figure 6 A step flow chart of step S105 provided for the embodiment of the present application;

[0061] Figure 7 Another step flow chart of the order quantity prediction method provided for the embodiment of the present application;

[0062] Figure 8 The overall flow chart of the order quantity prediction method provided for the embodiment of the present application;

[0063] Figure 9 The structure schematic diagram of the order quantity prediction device provided for the embodiment of the present application;

[0064] Figure 10 The hardware structure schematic diagram of the electronic device provided for the embodiment of the present application;

[0065] Figure 11 The structure schematic diagram of the storage medium provided for the embodiment of the present application. DETAILED DESCRIPTION

[0066] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application. It should be noted that although the functional modules are divided in the system schematic diagram, and the logical order is shown in the flow chart, in some cases, the steps shown or described can be executed in a different order from the module division in the system schematic diagram or the order in the flow chart. For the step numbers in the following embodiments, they are only set for the convenience of explanation and description, and the order between the steps is not limited in any way, and the execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0067] In the description of the present application, the meaning of multiple is two or more, and if the first, the second is described only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or implicitly indicating the sequence of indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0068] The order quantity prediction method provided by the embodiments of the present application can be applied to a terminal, can be applied to a server end, and can also be software running in the terminal or the server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a set-top box, etc.; the server end can be configured as a separate physical server, can be configured as a server cluster or a distributed system formed by multiple physical servers, can also be configured as a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform; and the software can be an application that implements the order quantity prediction method, but is not limited to the above forms.

[0069] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as a program module. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0070] It should be noted that in each specific embodiment of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the user, such as user information, user behavior data, user history data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards in relevant countries and regions. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for enabling the embodiments of the present application to normally operate will be obtained.

[0071] As Figure 1 Figure 1 shows a step flowchart of an order quantity prediction method provided by an embodiment of the present application, and Figure 1 The embodiments of the present application provide an order quantity prediction method, which specifically includes the following steps:

[0072] S101, acquire historical order data, advertisement promotion data, social media data and real-time exogenous variables, and obtain multi-source time series data through data alignment and data cleaning, perform feature extraction on the multi-source time series data to obtain a multi-modal feature matrix;

[0073] S102, perform multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network (MSCAN) to obtain an initial order quantity prediction sequence and an attention weight matrix;

[0074] S103, perform variational dynamic hybrid modeling (VDMHM) and variational inference optimization on the initial order quantity prediction sequence to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters;

[0075] S104, perform adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters based on a dynamic adaptive embedding mechanism (DAEL) to obtain refined feature embeddings;

[0076] S105, perform adaptive causal intervention analysis (ACIA) on the attention weight matrix, the optimized order quantity prediction sequence and the refined feature embeddings to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable.

[0077] The embodiment of the application fuses multi-modal data such as historical orders, advertisement promotion, social media dynamics and exogenous variables, realizes multi-level feature extraction and order quantity prediction through multi-scale causal attention network fusion wavelet transform and causal mask, obtains an optimized order quantity prediction sequence by quantifying the uncertainty of the prediction through variational dynamic hybrid modeling and variational inference, estimates the intervention effect of each intervention variable on the order quantity through adversarial reinforcement learning and adaptive causal intervention analysis, and finally obtains a target order quantity prediction sequence, thereby improving the comprehensiveness and accuracy of product order quantity prediction and providing strong support for marketing strategies, warehouse allocation and other decisions, which can effectively help enterprises maintain a competitive advantage in a complex and changing market environment.

[0078] As Figure 2 shown is a step flowchart of step S101 provided by the embodiment of the application, with reference to Figure 2 , as a further optional implementation, the multi-source time series data is obtained through data alignment and data cleaning, and the multi-modal feature matrix is obtained by performing feature extraction on the multi-source time series data, which specifically includes:

[0079] S1011, perform dynamic time warping on the historical order data, the advertisement promotion data, the social media data and the real-time exogenous variables to align the time, and perform missing value filling and outlier removal to obtain the multi-source time series data;

[0080] S1012, perform kernel principal component analysis on the multi-source time sequence data to extract a multi-modal feature matrix.

[0081] Specifically, the data collection and processing module is the cornerstone of the entire prediction system, responsible for obtaining data from multiple sources (such as historical orders, advertising, social media, and exogenous variables), and performing dynamic feature extraction and anomaly cleaning.

[0082] The present scheme designs a multi-level, multi-modal data pipeline that comprehensively covers key factors affecting order volume. The first type of data is historical order data, extracted from a big data lake, CRM system or data warehouse, represented by Y(t) as the order volume at time t, supporting flexible time granularity (such as hours, days, weeks, months). To capture short-term fluctuations and long-term trends in orders, Y(t) is subjected to multi-scale decomposition (such as through empirical mode decomposition EMD), separating noise, periodicity and trend components. The second type is advertising promotion data, represented by M(t) as the advertising volume or cost, and refined into cross-channel features M1(t), M2(t) etc. (such as live app, short video platform, search engine, social advertising), with the addition of dynamic material features (such as video length, creative type, CTR estimate). To improve data resolution, time-based exposure, click-through rate and conversion rate are extracted from advertising platform APIs, and their distribution characteristics (such as mean, variance, quantile) are calculated through sliding window statistics. The third type is social media data, obtained through the microblog open platform API or third-party crawling tools to obtain keyword mention volume S1(t), sentiment score S2(t) and topic burst index S3(t). In specific implementation, pre-trained language models (such as BERT or RoBERTa) are used to extract post semantic vectors, combined with dynamic topic models (such as Online LDA or neural topic model NTM) to cluster topic trends in real time, while sentiment analysis tools (such as enhanced VADER) are used to calculate positive and negative emotion proportions. To deal with fake volume or repeated content interference, a VAE-based anomaly detection module is introduced to train a reconstruction error threshold, dynamically filter abnormal data points and adjust feature weights (such as reducing the contribution of repeated content). The fourth type is exogenous variables, including competitor dynamics C(t) (such as competitor prices, promotion intensity, inventory changes), seasonal and holiday factors H(t) (such as Double 11, Spring Festival) and macroeconomic indicators (such as CPI, consumer confidence index). In addition, we add real-time multi-modal data such as user behavior trajectory B(t) (browsing depth, dwell time, add-to-cart rate) and environmental variables E(t) (weather index, sudden events or policy changes), obtained through IoT data or news APIs.

[0083] To improve data quality, the scheme designs a multi-source data alignment and cleaning process: first, use dynamic time warping (DTW) to align time series from different sources, solve the problem of inconsistent sampling frequency, and avoid the deviation caused by simple linear interpolation; then for missing values, combined with business context (such as missing orders on promotion day) to choose KNN-based interpolation or local Bayesian filling; for outliers, identify them by combining Isolation Forest and event labeling, retain reasonable fluctuations such as special date surges, and exclude data entry errors. After standardizing the multi-modal features, use Kernel PCA or UMAP dimensionality reduction to retain non-linear correlations and reduce redundancy, and finally form a structured feature matrix (i.e. multi-modal feature matrix) containing T(t), M(t), S(t), C(t), H(t), B(t), E(t) for subsequent module input.

[0084] As Figure 3 shown is a step flowchart of step S102 provided by the embodiment of the application, referring to Figure 3 , further as an optional implementation, the multi-modal feature matrix is subjected to multi-scale wavelet decomposition and causal attention calculation through a multi-scale causal attention network to obtain an initial order quantity prediction sequence and an attention weight matrix, which specifically includes:

[0085] S1021, decompose the multi-modal feature matrix into short-period components, medium-period components and long-period components through discrete wavelet transformation;

[0086] S1022, calculate the hidden state and information transmission of the short-period components, medium-period components and long-period components through the gate recurrent unit to obtain short-period prediction sequences, medium-period prediction sequences and long-period prediction sequences;

[0087] S1023, calculate the attention weights of the short-period components, medium-period components and long-period components based on the causal attention mechanism to obtain the attention weight matrix;

[0088] S1024, dynamically weight and fuse the short-period prediction sequences, medium-period prediction sequences and long-period prediction sequences according to the attention weight matrix to obtain the initial order quantity prediction sequence.

[0089] Specifically, after obtaining the structured feature matrix, preliminary time series prediction is performed, since the traditional attention mechanism (such as Transformer) has two problems in time series prediction: one is the lack of explicit causal constraints, which is easy to introduce future information leakage; the other is the insufficient modeling of dynamic features of multiple time scales. Therefore, the scheme designs a multi-scale causal attention network (MSCAN) to realize multi-level feature extraction and prediction through wavelet transformation and causal mask.

[0090] First, input the feature matrix X(t) = [M(t), S(t), C(t), H(t), B(t), E(t)] and decompose it into three time-scale components by Discrete Wavelet Transform (DWT):

[0091] X short (t): short-period component (e.g., intra-day fluctuations), capturing the instantaneous effect of ad clicks or social media post hotspots;

[0092] X mid (t): medium-period component (e.g., weekly trends), reflecting the cumulative impact of promotional activities or weekend effects;

[0093] X long (t): long-period component (e.g., monthly or seasonal), modeling macro trends and holiday effects.

[0094] The decomposition formula is:

[0095] X(t) = DWT(X(t)) = X short (t) + X mid (t) + X long (t)

[0096] where the Daubechies wavelet basis (e.g., db4) is used to ensure signal smoothness and minimize boundary effects.

[0097] At each scale, the proposed design improves the causal attention mechanism, introducing an adaptive time lag window and sparse attention to enhance computational efficiency. The attention calculation formula is:

[0098]

[0099] where: X, X, and X are the query, key, and value matrices, respectively, M and M are trainable parameters; M causal is the causal mask matrix, ensuring that only data before t is considered at time t, and is a lower triangular matrix; τ max is the adaptive time lag window, dynamically adjusted by feature mutual information (e.g., 3 days for ad effects and 1 day for social media); d k is the dimension of the key vector, used to scale to avoid gradient vanishing.

[0100] To reduce computational complexity, we introduce sparse attention based on feature importance (e.g., SHAP values) to select the Top-K relevant time steps, reducing the time complexity from O(T 2 ) to O(TlogT). The multi-scale output is fused by dynamic weighting:

[0101]

[0102] where W short, W mid, W long Adaptive learning through gating mechanism (such as Gated Linear Unit, GLU) avoids manual parameter tuning. After obtaining the preliminary prediction result through the multi-scale causal attention network, the result is input to the next step for uncertainty modeling.

[0103] As Figure 4 Figure 2 shows a step flowchart of step S103 according to an embodiment of the present application. As shown in Figure 4 As a further optional implementation, the initial order quantity prediction sequence is subjected to variational dynamic mixture modeling and variational inference optimization to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters, which specifically include:

[0104] S1031, constructing an initial variational distribution according to the initial order quantity sequence, and inputting the initial variational distribution into an uncertainty network to obtain an order quantity mean and an order quantity variance;

[0105] S1032, modeling the initial order quantity sequence, the order quantity mean and the order quantity variance according to a Gaussian mixture model to obtain a dynamic Gaussian mixture distribution;

[0106] S1033, minimizing the KL divergence loss through variational inference based on the approximate posterior distribution to obtain an optimized dynamic Gaussian mixture distribution;

[0107] S1034, determining an optimized order quantity prediction sequence, an optimized order quantity mean and an optimized order quantity variance corresponding to the optimized order quantity prediction sequence according to the optimized dynamic Gaussian mixture distribution, and generating a hidden variable distribution parameter according to the optimized order quantity mean and the optimized order quantity variance.

[0108] Specifically, after obtaining the preliminary order quantity prediction of MSCAN, the actual order quantity Y(t) often presents a multi-peak distribution (such as a sharp increase in orders on promotion days and a smooth distribution at other times), and it is difficult to accurately model using a traditional single Gaussian assumption (such as LSTM+MAE). Therefore, the present scheme proposes variational dynamic mixture modeling (VDMHM) combined with a Gaussian mixture model (GMM) and variational inference to quantify the uncertainty of the prediction. Let Y(t) follow a dynamic Gaussian mixture distribution:

[0109]

[0110] where: M is the number of mixture components (adaptively determined through Dirichlet process to avoid presetting); π m (t) is the dynamic weight of the mth component, satisfying μm (t) and σ m (t) are mean and standard deviation predicted by the neural network.

[0111] In a specific implementation, the output of MSCAN As backbone features, input to an auxiliary uncertainty network (UN), which contains three sub-networks:

[0112] Mean network:

[0113] Variance network: (softplus ensures positive values);

[0114] Weight network:

[0115] To optimize the parameters, we use variational inference, introducing an approximate posterior distribution q(z(t) | X(t)) (latent variable z(t) represents the mixture component selection), the goal is to maximize the evidence lower bound (ELBO):

[0116] ELBO = E q(z(t)) [log p(Y(t), z(t) | X(t))] - KL(q(z(t) | X(t)) || p(z(t)))

[0117] Where: p(Y(t), z(t) | X(t)) is the joint likelihood; KL is the Kullback-Leibler divergence, which measures the difference between the approximate distribution and the prior.

[0118] The final prediction is the expected value:

[0119]

[0120] The confidence interval is generated by Monte Carlo sampling (such as the 95% interval [μ m (t) ± 1.96 σ m (t)] weighted combination).

[0121] As shown in Figure 5 is a step flowchart of step S104 provided by the embodiment of the application, referring to Figure 5 , further as an optional implementation, based on the dynamic adaptive embedding mechanism, the hidden variable distribution parameter is subjected to adversarial reinforcement learning and feature reconstruction to obtain refined feature embedding, which specifically includes:

[0122] S1041, generating a dynamic pseudo-distribution parameter according to the historical context of the hidden variable distribution parameter through a time-conditioned generator, and judging the distribution consistency loss of the dynamic pseudo-distribution parameter and the hidden variable distribution parameter through a discriminator;

[0123] S1042, determine a dynamic pseudo-order quantity prediction sequence according to the dynamic pseudo-distribution parameter, and determine a prediction consistency loss of the dynamic pseudo-order quantity prediction sequence and the optimized order quantity prediction sequence;

[0124] S1043, determine a joint loss of the adversarial training according to the distribution consistency loss and the prediction consistency loss, and optimize the joint loss through deep reinforcement learning to obtain a trained time sequence condition generator;

[0125] S1044, generate a target distribution parameter through the trained time sequence condition generator, and perform feature reconstruction on the target distribution parameter to obtain a refined feature embedding.

[0126] Specifically, to enhance the robustness of the model to abnormal scenarios (such as competitor price reduction, real-time and emotion), dynamic adversarial enhanced learning (DAEL) is introduced, combining the ideas of generative adversarial network (GAN) and reinforcement learning (RL). Traditional GAN generates static pseudo-data, which is difficult to adapt to time sequence dynamics. Therefore, a temporal conditional generator (TCG) is designed in the scheme:

[0127] X fake (t)=G(X(t-τ:t),z(t))

[0128] Where z(t) is a noise vector, G is a generator (such as a sequence generation network based on GRU), and the conditional input X(t-τ:t) provides historical context.

[0129] The discriminator D judges the distribution consistency of X fake (t) and the real data X real (t):

[0130] L adv =-logD(X real (t))-log(1-D(X fake (t)))

[0131] The generator aims to minimize the discriminator's ability to distinguish, while adding a prediction consistency constraint:

[0132] L G =L adv +ηL consist (Y pred (X real ),Y pred (X fake ))

[0133] Where L consist is the prediction consistency loss (such as MSE), and η is a balance coefficient.

[0134] To achieve dynamic adjustment of adversarial strength, we introduce a reinforcement learning (RL) module based on deep Q-network. The core goal of this module is to balance the prediction accuracy of the main prediction model on normal data and the robustness on adversarial generated data (pursue high prediction consistency). The state S(t) of RL is designed as a vector containing key information at the current time.

[0135] In specific implementation, the original high-dimensional input X(t) is extracted into low-dimensional features f X (t) through dimension reduction, the prediction result Y real (t) of the main model on the current real data X pred (t), and the uncertainty σ(t) of the model prediction are spliced. That is, S(t) = concat(f X (t), Y pred (t), σ(t)). Since the state space may contain continuous values and has high dimension, we do not use the traditional table type Q-learning, but use a neural network Q(S, A; θ) to approximate the Q function.

[0136] The action A(t) of RL corresponds to the weight λ(t) of the adversarial loss. In order to adapt to the standard DQN framework, we discretize the continuous λ(t) value into a predefined, finite action set, for example A = {λ1, λ2, …, λ k}, which may contain values such as A = {0.0, 0.1, 0.5, 1.0, 2.0}. A value of 0 means temporarily turning off adversarial training, and a larger value means enhancing adversarial resistance. The task of RL is to select an optimal λ i as A(t) from this set at each decision time step t. Ensure that the model is balanced between normal and abnormal scenarios.

[0137] As Figure 6 shown is a step flowchart of step S105 provided by the embodiment of the application, with reference to Figure 6 , further as an optional implementation, the adaptive causal intervention analysis is performed on the attention weight matrix, the optimized order quantity prediction sequence and the refined feature embedding, to obtain a target order quantity prediction sequence and a causal effect estimation value of each intervention variable, which specifically includes:

[0138] S1051, constructing a dynamic causal graph with the time series features of each modality of the multi-modal feature matrix as nodes, and updating the edge weight between nodes through Granger causality test and attention weight matrix;

[0139] S1052, calculating the causal effect estimation value of each intervention variable through counterfactual reasoning according to the refined feature embedding and the dynamic causal graph;

[0140] S1053, updating the optimized order quantity prediction sequence according to the causal effect estimation value, to obtain a target order quantity prediction sequence.

[0141] Specifically, to provide business-operable insights, an adaptive causal intervention analysis (ACIA) is designed to quantify the causal effect of advertising or social media based on Do-Calculus and dynamic Bayesian networks (DBN). Traditional causal analysis (such as difference-in-difference) ignores time dependence, and ACIA models intervention effects through a time-series causal graph:

[0142] ΔY(t)=E[Y(t)∣do(M(t)=m1)]-E[Y(t)∣do(M(t)=m0)]

[0143] where do(M(t)=m1) represents an intervention (such as a 10% increase in advertising budget).

[0144] In a specific implementation, a dynamic causal graph G(t)=(V(t),E(t)) is constructed: V(t)=[Y(t),M(t),S(t),C(t),…] is a node; E(t) is an edge, which is dynamically updated through Granger causality test and attention weight.

[0145] Intervention effects are calculated through counterfactual prediction:

[0146] Y cf (t)=f MSCAN (X(t)∣M(t)=m1,z(t))

[0147] where z(t) is sampled from VDMHM to ensure uncertainty consistency. To adaptively adjust causal assumptions, an online verification module is added: based on A / B test results or synthetic control method, the edge weights of G(t) are regularly updated.

[0148] ACIA combines dynamic causal graphs and counterfactual reasoning to provide time-dependent intervention analysis, and adaptive verification ensures the accuracy of causal explanations.

[0149] As Figure 7 shown is another step flowchart of the order quantity prediction method provided by an embodiment of the application, with reference to Figure 7 , further as an optional implementation, the order quantity prediction method further includes the following steps:

[0150] S201, obtaining an order quantity real sequence of a target period, and determining a real order quantity mean and a real order quantity variance according to the order quantity real sequence;

[0151] S202, calculating an order quantity prediction loss according to the order quantity real sequence and the target order quantity prediction sequence;

[0152] S203, calculating the evidence lower bound loss according to the order quantity real sequence and the optimized dynamic Gaussian mixture distribution;

[0153] S204, calculating the real distribution loss according to the real order quantity mean, the real order quantity variance and the hidden variable distribution parameter;

[0154] S205, calculating the causal prediction loss according to the change rate of the order quantity real sequence at each time point and the change rate of the target order quantity prediction sequence at each time point;

[0155] S206, determining the overall loss value according to the order quantity prediction loss, the evidence lower bound loss, the real distribution loss and the causal prediction loss;

[0156] S207, updating the parameters of the multi-scale causal attention network, the variational dynamic mixture modeling, the adversarial reinforcement learning and the adaptive causal intervention analysis according to the overall loss value.

[0157] Specifically, after obtaining the target order quantity prediction sequence Y pred (t) of the target period, the real order quantity is recorded until the target period to obtain the order quantity real sequence Y true (t). The overall loss function can be determined in combination with the above steps as follows:

[0158] L total = L pred + αL unc + βL adv + γL causal

[0159] Wherein:

[0160] The order quantity prediction loss

[0161] The evidence lower bound loss L unc = -ELBO;

[0162] The real distribution loss L adv = -logD(X real )-log(1-D(X fake ));

[0163] The causal prediction loss L causal =∑ t (ΔY pred (t)-ΔY true (t)) 2 .

[0164] The overall loss value of the prediction is calculated based on the overall loss function, and then the parameters of the multi-scale causal attention network, the variational dynamic hybrid modeling, the adversarial reinforcement learning, and the adaptive causal intervention analysis are updated according to the overall loss value, so that the order quantity prediction of the next time is more accurate.

[0165] As Figure 8 shown is a whole process schematic diagram of an order quantity prediction method provided by an embodiment of the application, taking order quantity prediction during Double 11 as an example, the daily order quantity is predicted during the Double 11 promotion activity from November 1, 2025 to November 11, 2025, so as to optimize the inventory allocation and advertisement placement strategy. The system needs to integrate historical order data, advertisement placement records, social media dynamics and exogenous variables (such as weather, holiday effect), and output the prediction results and causal analysis.

[0166] First step: data collection and preprocessing

[0167] Logical relationship: as the starting point of the whole process, provide the input data required for subsequent modeling.

[0168] Executing body: data collection and processing module (distributed data pipeline, running on cloud server).

[0169] Trigger condition: start at the activity preparation stage (0 o'clock on October 31, 2025), and collect the data up to the current time in real time.

[0170] Processing object: multi-source data, including historical order data (daily order quantity during Double 11 in the past two years), advertisement placement data (placement logs of electronic channels such as short video platforms, WeChat public accounts and search engines), social media data (Double 11 related posts on Weibo), and exogenous variables (weather forecast, holiday calendar).

[0171] Specific actions:

[0172] i. Extract order data Y from the e-commerce background database t (format: CSV, field contains timestamp order quantity).

[0173] ii. Get placement data A through the advertisement platform API t (JSON, contains placement time, channel, cost, click rate).

[0174] iii. Use the Weibo open platform API to crawl social media posts, extract keyword mention volume S t and sentiment score (calculated by RoBERTa model).

[0175] iv. Get weather index and holiday factor E t (API returns structured data).

[0176] v. Align time series (using DTW algorithm), impute missing values (KNN interpolation), remove outliers (Isolation Forest).

[0177] vi. Standardize and reduce dimensionality (KernelPCA), generate feature tensor.

[0178] Result: Multimodal feature tensor X = [batch_size, time_steps, feature_dim] (e.g., [10, 24, 128] represents 10 days, 24 hours per day, 128-dimensional features).

[0179] Result usage: As input for MSCAN, for feature extraction and preliminary prediction.

[0180] Action: Ensure high quality and consistency of input data, provide foundation for subsequent modeling.

[0181] Second step: Multiscale causal attention modeling (MSCAN)

[0182] Logical relationship: Based on the feature tensor from the first step, extract multiscale temporal features and identify causal relationships.

[0183] Executing body: MSCAN module (runs on GPU cluster, such as NVIDIA A100).

[0184] Trigger condition: Start immediately after the first step completes feature tensor generation (data processing completed on October 31).

[0185] Processing object: Feature tensor X output by the first step.

[0186] Specific actions:

[0187] i. Use discrete wavelet transform (DWT) to decompose X into short-period X s , medium-period X m , and long-period X l components.

[0188] ii. Apply causal attention mechanism to each component (set causal mask, ensure that t only depends on t previous data).

[0189] iii. Calculate attention weights where Q t and K i are query and key vectors.

[0190] iv. Weighted fusion of multiscale output: Y pred (t) = w s Y s (t) + w m Y m(t) + w l Y l (t).

[0191] Result: Preliminary prediction sequence Y pred = [batch_size, time_steps, feature_dim] (e.g., [10, 24, 1] means 10 days of hourly prediction values) and attention weight matrix A = [batch_size, time_steps, feature_dim].

[0192] Result usage: Input to VDMHM module, further optimize prediction and uncertainty modeling.

[0193] Action: Capture multi-time scale order fluctuations and preliminarily quantify the importance of features (e.g., ad spend).

[0194] Step 3: Variational Dynamic Multi-Modal Hybrid Modeling (VDMHM)

[0195] Logical relationship: Based on the preliminary prediction of the second step, enhance uncertainty estimation and dynamic correlation modeling.

[0196] Executing body: VDMHM module (runs on a same-GPU cluster).

[0197] Trigger: Triggered after the second step generates a preliminary prediction sequence (MSCAN computation completed on October 31).

[0198] Processing object: MSCAN's prediction sequence Y pred and original feature tensor X.

[0199] Specific actions:

[0200] i. Construct the variational distribution q(z t | Z, Y pred ), estimate the mean μ t and variance

[0201] ii. Use a hybrid model to fuse multi-modal data:

[0202] iii. Optimize the KL divergence loss: L KL = D KL (q(z t || p(z t )).

[0203] iv. Output the optimized prediction sequence and hidden variable distribution.

[0204] Result: Optimized prediction sequence Y opt= [batch_size, time_steps, 1] and latent variable distribution parameters Z = [batch_size, time_steps, latent_dim] (e.g., [10, 24, 16]).

[0205] Result usage: Input to DAEL module for feature refinement and noise handling.

[0206] Action: Improve the robustness of predictions and quantify uncertainty, providing support for subsequent causal analysis.

[0207] Step 4: Dynamic Adaptive Embedding Learning (DAEL)

[0208] Logical relationship: Based on the output of Step 3, refine features and enhance the model's adaptability to noise.

[0209] Executing body: DAEL module (runs on GPU cluster).

[0210] Trigger condition: Triggered after completing the calculation of latent variable distribution in Step 3 (VDMHM output generation on October 31).

[0211] Processing object: Latent variable distribution Z and original feature tensor X of VDMHM.

[0212] Specific actions:

[0213] i. Construct an adversarial network to generate dynamic pseudo data X fake (through the conditional generator).

[0214] ii. Calculate embedding representation: E t = f(X t , z t ; θ), where f is an adaptive neural network.

[0215] iii. Optimize adversarial loss:

[0216] Result: Refined feature embedding E = [batch_size, time_steps, refined_dim] (e.g., [10, 24, 64]) and reconstruction error.

[0217] Result usage: Input to ACIA module for causal effect analysis.

[0218] Action: Improve the quality of feature representation and enhance the model's generalization ability to abnormal data (such as the surge in orders during Double 11).

[0219] Step 5: Adaptive Causal Intervention Analysis (ACIA)

[0220] Logical relationship: integrate the previous outputs, generate the final prediction and analyze the causal effect.

[0221] Execution subject: ACLA module (running on a cloud server, combining CPU and GPU).

[0222] Trigger condition: triggered after feature refinement in the fourth step (DAEL output generated on October 31st).

[0223] Specific actions:

[0224] i. Construct a dynamic causal graph G = (V, E) and update the edge weight through Granger causality test.

[0225] ii. Calculate the intervention effect:

[0226] iii. Generate the final prediction sequence and attach the confidence interval (based on Z sampling).

[0227] Results: final output, including prediction sequence Y final (orders from November 1st to 11th), 95% confidence interval, causal effect estimate (such as the average effect of increasing advertising by 10%).

[0228] Results usage: output in JSON format for enterprises to adjust inventory and advertising strategies.

[0229] Effect: provide high-precision prediction and actionable business insights.

[0230] The method flow of the embodiment of the application is described above. It can be understood that the embodiment of the application integrates historical orders, advertising promotion, social media dynamics and exogenous variables and other multi-modal data, realizes multi-level feature extraction and order quantity prediction through multi-scale causal attention network fusion wavelet transform and causal mask, optimizes the order quantity prediction sequence through variational dynamic hybrid modeling and variational inference to quantify the uncertainty of prediction, estimates the intervention effect of each intervention variable on the order quantity through adversarial reinforcement learning and adaptive causal intervention analysis, and finally obtains the target order quantity prediction sequence, improves the comprehensiveness and accuracy of product order quantity prediction, provides strong support for marketing strategy, warehouse allocation and other decisions, and can effectively help enterprises maintain a competitive advantage in complex and changing market environment.

[0231] As Figure 9 shown is a structural schematic diagram of an order quantity prediction device provided by the embodiment of the application, with reference to Figure 8 , the embodiment of the application provides an order quantity prediction device, which comprises:

[0232] The data acquisition and processing module is configured to acquire historical order data, advertising promotion data, social media data, and real-time exogenous variables, and obtain multi-source time series data through data alignment and data cleaning, and perform feature extraction on the multi-source time series data to obtain a multi-modal feature matrix.

[0233] The multi-scale causal attention module is configured to perform multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network to obtain an initial order quantity prediction sequence and an attention weight matrix.

[0234] The variational dynamic hybrid modeling module is configured to perform variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters.

[0235] The dynamic adaptive embedding learning module is configured to perform adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters based on a dynamic adaptive embedding mechanism to obtain refined feature embeddings.

[0236] The adaptive causal intervention analysis module is configured to perform adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence, and the refined feature embeddings to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable.

[0237] Specifically, the entire order quantity prediction device is composed of multiple modules, including a data acquisition and processing module, a multi-scale causal attention module, a variational dynamic hybrid modeling module, a dynamic adaptive embedding learning module, and an adaptive causal intervention analysis module. These modules cooperate through precise data flow to form a prediction system. The relationship between each module and the specific technical features of the input and output interfaces are described in detail below.

[0238] 1. Data acquisition and processing module: The data acquisition and processing module is responsible for extracting features from multi-source heterogeneous data (such as order records, advertising data, social media dynamics, and external market indicators), and performing cleaning, standardization, and multi-modal alignment to provide high-quality input data for subsequent modules.

[0239] Input:

[0240] 1) Raw data: time series order data (such as CSV format, containing timestamp and order quantity), advertising log (JSON format, containing time, channel, and cost), social media posts (unstructured data such as text and pictures).

[0241] 2) External signals: weather data, holiday calendar, competitor activity data (structured data returned by API call).

[0242] Output:

[0243] 1) Processed multi-modal feature tensor [batch_size, time_steps, feature_dim] containing numerical features (e.g., order volume, ad spend) and embedded features (e.g., social media sentiment vectors).

[0244] 2) Time window metadata: Time window information, feature normalization parameters (e.g., mean and variance).

[0245] 2. Multi-scale causal attention module: Receives multi-modal feature tensor from data acquisition and processing module, utilizes multi-scale attention mechanism to capture short-term fluctuations and long-term trends in time series, and identifies key driving factors (e.g., direct impact of ad placement on orders) through causal attention. Its output is preliminary prediction results and attention weights for subsequent optimization.

[0246] Input:

[0247] 1) Multi-modal feature tensor [batch_size, time_steps, feature_dim] from data acquisition module.

[0248] 2) Time window metadata.

[0249] Output:

[0250] 1) Preliminary prediction sequence: [batch_size, time_steps, 1], representing order volume prediction at each time step.

[0251] 2) Attention weight matrix: [batch_size, time_steps, feature_dim], representing the contribution of each feature to the prediction.

[0252] 3. Variational dynamic hybrid modeling module: Obtains preliminary prediction sequence and attention weights from multi-scale causal attention module, further models the uncertainty and dynamic correlation of multi-modal data (e.g., potential relationship between social media sentiment and order volume). Its variational inference capability enhances the robustness of the system and provides more refined latent variable representation for DAEL.

[0253] Input:

[0254] 1) Preliminary prediction sequence [batch_size, time_steps, 1] from MSCAN.

[0255] 2) Multi-modal feature tensor [batch_size, time_steps, feature_dim] from data acquisition module.

[0256] Output:

[0257] 1) Optimized prediction sequence: [batch_size, time_steps, 1].

[0258] 2) Hidden variable distribution parameters: mean and variance [batch_size, time_steps, latent_dim].

[0259] 4. Dynamic Adaptive Embedding Learning Module Function and Connection: Obtain hidden variable distribution and optimized prediction sequence from Variational Dynamic Hybrid Modeling Module, capture nonlinear patterns and noise (such as abnormal order fluctuations) in data through dynamic adaptive embedding mechanism. Its output is a refined feature representation for ACIA to conduct causal analysis.

[0260] Input:

[0261] 1) Hidden variable distribution parameters of VDMHM [batch_size, time_steps, latent_dim].

[0262] 2) Original feature tensor of data acquisition module [batch_size, time_steps, feature_dim].

[0263] Output:

[0264] 1) Refined feature embedding: [batch_size, time_steps, refined_dim].

[0265] 2) Reconstruction error: auxiliary signal for supervised training.

[0266] 5. Adaptive Causal Intervention Analysis Module Function and Connection: Integrate the outputs of the aforementioned modules, use causal intervention analysis to verify the causal explanatory power of the prediction (such as "whether increasing advertising investment significantly increases order volume"), and generate the final prediction result. Its results are fed back to the comprehensive loss function to optimize the entire system.

[0267] Input:

[0268] 1) Refined feature embedding of DAEL [batch_size, time_steps, refined_dim].

[0269] 2) Optimized prediction sequence of VDMHM [batch_size, time_steps, 1].

[0270] 3) Attention weight matrix of MSCAN [batch_size, time_steps, feature_dim].

[0271] Output:

[0272] 1) Final prediction sequence: the order volume prediction sequence of future time steps, the prediction uncertainty of 95% confidence interval, and the contribution of each input feature to the prediction.

[0273] 2) Causal effect estimation: the average causal effect (ATE) of an intervention variable (such as advertising spending).

[0274] The contents in the above method embodiments are all applicable to the device embodiments, the device embodiments specifically implement the functions of the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0275] The embodiment of the present application also provides an electronic device, which comprises a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing the connection communication between the processor and the memory, and the program is executed by the processor to realize the order volume prediction method. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0276] As shown in Figure 10 The hardware structure schematic diagram of the electronic device provided by the embodiment of the present application is shown in the figure, and the embodiment of the present application provides an electronic device, which comprises: Figure 10

[0277] The processor 1001 can be implemented in the form of a general CPU (Central Processing Unit, central processor), a microprocessor, an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), or one or more integrated circuits, and is used to execute related programs to realize the technical solutions provided by the embodiment of the present application.

[0278] The memory 1002 can be implemented in the form of a ROM (Read Only Memory, read-only memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory, random access memory). The memory 1002 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present application are realized by software or firmware, the related program codes are saved in the memory 1002 and are called and executed by the processor 1001 to realize the order volume prediction method of the embodiment of the present application.

[0279] The input / output interface 1003 is used to realize information input and output.

[0280] ​The communication interface 1004 is configured to realize the communication interaction between the device and other devices, and can realize the communication through a wired manner (for example, a USB, a network cable and the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth and the like).

[0281] The bus 1005 is configured to transmit information between various components (for example, the processor 1001, the memory 1002, the input / output interface 1003 and the communication interface 1004) of the device.

[0282] The processor 1001, the memory 1002, the input / output interface 1003 and the communication interface 1004 are connected to each other through the bus 1005.

[0283] As shown in Figure 11 FIG. 1 is a structural schematic diagram of a storage medium provided by an embodiment of the present application, and Figure 11 The embodiment of the present application further provides a storage medium, which is a computer readable storage medium, is used for computer readable storage, and stores one or more programs 1101, which can be executed by one or more processors to realize the order quantity prediction method.

[0284] The memory is a non-transient computer readable storage medium, and can be used to store a non-transient software program and a non-transient computer executable program. In addition, the memory can include a high-speed random access memory, and can further include a non-transient memory, for example, at least one magnetic disk storage device, a flash memory device or other non-transient solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and the remote memory can be connected to the processor through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0285] The embodiment of the present application further discloses a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device can read the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to execute Figure 1 the method shown in FIG. 1.

[0286] In alternative embodiments, the functions / operations in the flow diagrams can occur in sequences other than those depicted. For example, two operations shown in succession can in fact be executed substantially concurrently or the operations can sometimes be executed in the reverse order depending upon the functionality / operations involved. Such variations are contemplated to be within the scope of the present application. Embodiments presented and described in the flow diagrams are examples only and are used to provide an overall understanding of the method of the application. The disclosed methods are not limited to the operations and logical flows presented in this application. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations of a larger operation are performed in parallel rather than sequentially. The scope of the application is defined by the appended claims.

[0287] Furthermore, although the present application has been described in the context of functional modules, it is to be understood that one or more of the functions and / or features described above can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It will also be appreciated that detailed discussion of the actual implementation of each module is not necessary for an understanding of the application. Rather, the properties, functions and internal relationships of the various functional modules disclosed in the devices herein are deemed to be of a character as will be readily understood by those skilled in the art. Accordingly, the actual implementation of the modules is within the routine skill of engineers familiar with the properties, functions and internal relationships of the modules, given the disclosure herein of the attributes of the various functional modules. Therefore, the application illustrated in the claims is not limited to the embodiments described hereinabove by way of example. Rather, the scope of the application is to be determined by the full breadth of the claims and their equivalents.

[0288] If the above functions are implemented in software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the technical solutions that make essential contributions to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described above in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0289] The logic and / or steps represented in the flow diagrams or otherwise described herein, for example, can be embodied in non-transitory computer- readable media, executed by an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions, or in conjunction with which the instructions can be executed. In the context of this specification, a "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium.

[0290] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and stored in a computer memory.

[0291] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), and / or the like.

[0292] In the above description of the present specification, reference has been made to descriptive terms such as "one embodiment / exemplification", "another embodiment / exemplification", or "some embodiments / exemplifications" etc. It is understood that such terms are not intended to mean that the described specific feature, structure, material or characteristic was included in only one embodiment / exemplification. The illustrative descriptions are not intended to limit the scope of the present specification to the described embodiments / exemplifications. Furthermore, the described specific features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments / exemplifications.

[0293] While the embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary and are not to be construed as limiting the scope of the application. The scope of the application is defined by the appended claims and their equivalents.

[0294] The above is the specific description of the preferred embodiment of the application, but the application is not limited to the above-mentioned embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the application.

Claims

1. An order volume prediction method characterized by, The method comprises the following steps: obtaining historical order data, advertisement promotion data, social media data and real-time exogenous variables, and obtaining multi-source time series data through data alignment and data cleaning, and obtaining a multi-modal feature matrix through feature extraction on the multi-source time series data; performing multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network to obtain an initial order quantity prediction sequence and an attention weight matrix; performing variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence to obtain an optimized order quantity prediction sequence and hidden variable distribution parameters; performing adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters based on a dynamic adaptive embedding mechanism to obtain refined feature embeddings; performing adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence and the refined feature embeddings to obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable.

2. The order volume prediction method of claim 1, wherein, The multi-source time series data is obtained through data alignment and data cleaning, and the multi-modal feature matrix is obtained through feature extraction on the multi-source time series data, which specifically comprises: performing dynamic time warping on the historical order data, the advertisement promotion data, the social media data and the real-time exogenous variables to align the time, and performing missing value filling and outlier removal to obtain the multi-source time series data; performing kernel principal component analysis on the multi-source time series data to extract the multi-modal feature matrix.

3. The order volume prediction method of claim 1, wherein, The multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through the multi-scale causal attention network to obtain the initial order quantity prediction sequence and the attention weight matrix specifically comprises: decomposing the multi-modal feature matrix into short-period components, medium-period components and long-period components through discrete wavelet transformation; performing hidden state calculation and information transmission on the short-period components, the medium-period components and the long-period components through a gated recurrent unit to obtain short-period prediction sequences, medium-period prediction sequences and long-period prediction sequences; calculating the attention weights of the short-period components, the medium-period components and the long-period components based on a causal attention mechanism to obtain the attention weight matrix; performing dynamic weighted fusion on the short-period prediction sequences, the medium-period prediction sequences and the long-period prediction sequences according to the attention weight matrix to obtain the initial order quantity prediction sequence.

4. The order volume prediction method of claim 1, wherein, The variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence to obtain the optimized order quantity prediction sequence and the hidden variable distribution parameters specifically comprises: constructing an initial variational distribution according to the initial order quantity sequence, inputting the initial variational distribution into an uncertainty network to estimate an order quantity mean and an order quantity variance; modeling the initial order quantity sequence, the order quantity mean and the order quantity variance based on a Gaussian mixture model to obtain a dynamic Gaussian mixture distribution; minimizing the KL divergence loss through variational inference based on the approximate posterior distribution to obtain the optimized dynamic Gaussian mixture distribution; determining the optimized order quantity prediction sequence, the optimized order quantity mean and the optimized order quantity variance according to the optimized dynamic Gaussian mixture distribution, and generating the latent variable distribution parameter according to the optimized order quantity mean and the optimized order quantity variance.

5. The order volume prediction method of claim 4, wherein, The dynamic self-adaptive embedding mechanism is used to perform adversarial reinforcement learning and feature reconstruction on the latent variable distribution parameter to obtain refined feature embedding, and the refined feature embedding specifically includes: A dynamic pseudo-distribution parameter is generated by a time-conditioned generator according to the historical context of the latent variable distribution parameter, and a distribution consistency loss of the dynamic pseudo-distribution parameter and the latent variable distribution parameter is judged by a discriminator; A dynamic pseudo-order quantity prediction sequence is determined according to the dynamic pseudo-distribution parameter, and a prediction consistency loss of the dynamic pseudo-order quantity prediction sequence and the optimized order quantity prediction sequence is determined; A joint loss of adversarial training is determined according to the distribution consistency loss and the prediction consistency loss, and the joint loss is optimized by deep reinforcement learning to obtain a trained time-conditioned generator; A target distribution parameter is generated by the trained time-conditioned generator, and feature reconstruction is performed on the target distribution parameter to obtain the refined feature embedding.

6. The order volume prediction method of claim 1, wherein, The attention weight matrix, the optimized order quantity prediction sequence and the refined feature embedding are subjected to adaptive causal intervention analysis to obtain a target order quantity prediction sequence and a causal effect estimate value of each intervention variable, and the adaptive causal intervention analysis specifically includes: A dynamic causal graph is constructed by taking the time-series features of each modality of the multi-modal feature matrix as nodes, and the edge weights between nodes are updated by Granger causality test and the attention weight matrix; The causal effect estimate value of each intervention variable is calculated by counterfactual reasoning according to the refined feature embedding and the dynamic causal graph; The optimized order quantity prediction sequence is updated according to the causal effect estimate value to obtain the target order quantity prediction sequence.

7. The order volume prediction method of any one of claims 1 to 6, wherein, The order quantity prediction method further includes the following steps: An order quantity real sequence of a target period is obtained, and a real order quantity mean and a real order quantity variance are determined according to the order quantity real sequence; An order quantity prediction loss is calculated according to the order quantity real sequence and the target order quantity prediction sequence; An evidence lower bound loss is calculated according to the order quantity real sequence and the optimized dynamic Gaussian mixture distribution; A real distribution loss is calculated according to the real order quantity mean, the real order quantity variance and the latent variable distribution parameter; A causal prediction loss is calculated according to the change rate of the order quantity real sequence at each time point and the change rate of the target order quantity prediction sequence at each time point; A total loss value is determined according to the order quantity prediction loss, the evidence lower bound loss, the real distribution loss and the causal prediction loss; The parameters of the multi-scale causal attention network, the variational dynamic mixture modeling, the adversarial reinforcement learning and the adaptive causal intervention analysis are updated according to the total loss value.

8. An order amount prediction device characterized by comprising: ​ The data acquisition and processing module is configured to acquire historical order data, advertisement promotion data, social media data, and real-time exogenous variables, obtain multi-source time series data through data alignment and data cleaning, and obtain a multi-modal feature matrix through feature extraction on the multi-source time series data. The multi-scale causal attention module is configured to perform multi-scale wavelet decomposition and causal attention calculation on the multi-modal feature matrix through a multi-scale causal attention network, and obtain an initial order quantity prediction sequence and an attention weight matrix. The variational dynamic hybrid modeling module is configured to perform variational dynamic hybrid modeling and variational inference optimization on the initial order quantity prediction sequence, and obtain an optimized order quantity prediction sequence and hidden variable distribution parameters. The dynamic adaptive embedding learning module is configured to perform adversarial reinforcement learning and feature reconstruction on the hidden variable distribution parameters based on a dynamic adaptive embedding mechanism, and obtain refined feature embeddings. The adaptive causal intervention analysis module is configured to perform adaptive causal intervention analysis on the attention weight matrix, the optimized order quantity prediction sequence, and the refined feature embeddings, and obtain a target order quantity prediction sequence and causal effect estimates of each intervention variable.

9. An electronic device, comprising: The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and the program, when executed by the processor, realizes the steps of the order quantity prediction method according to any one of claims 1 to 7.

10. A storage medium, the storage medium being a computer-readable storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs executable by one or more processors to realize the steps of the order quantity prediction method according to any one of claims 1 to 7.