Method and system for predicting retail sales volume and computer readable storage medium
Through feature coding and attention interaction processing of multi-source heterogeneous data, the problem of insufficient data fusion and periodic features in retail sales forecasts is solved, and higher prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510772275.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-08
AI Technical Summary
Existing retail sales forecasting technologies are difficult to effectively integrate multi-source heterogeneous feature data, especially static attributes and dynamic timing information, and fail to fully utilize the cyclical features of retail business, resulting in insufficient prediction accuracy.
Feature encoding processing is used to convert multi-source heterogeneous feature data into a unified vector representation, and through the learningable global token vector and the attention interaction between the endogenous and exogenous variables, the multi-head hidden attention mechanism and a hybrid expert network are used to adapt to the retail business cycle characteristics and carry out in-depth information integration.
It improves the accuracy of retail sales forecasts, significantly reduces the average absolute percentage error, enhances the adaptability and scalability of the model, and is suitable for large-scale retail scenarios.
Smart Images

Figure CN120278758A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to machine learning, and more particularly to methods, systems, and computer-readable storage media for predicting retail sales volume. Background Art
[0002] Time series prediction plays a key role in many industries, especially in the retail field. Accurate sales volume prediction is crucial for inventory management, supply chain optimization, personnel scheduling, and marketing strategy formulation. To meet this demand, the industry has developed a variety of prediction techniques. These techniques are extensive, including both traditional statistical models such as autoregressive integrated moving average model (ARIMA) and its variants, and also deep learning-based methods that have emerged in recent years, such as recurrent neural network (RNN) and long short-term memory network (LSTM).
[0003] However, traditional prediction methods often face significant challenges when dealing with the complex, multi-source, and heterogeneous data commonly found in modern retail scenarios. Specifically, it is difficult for existing technologies to effectively integrate feature data of different natures. For example, it is difficult to combine relatively static attribute information (such as product brand, store location) with dynamically changing time series information (such as historical sales volume, promotion intensity, weather changes) for comprehensive analysis. In addition, some attempts to directly apply advanced models from other fields such as natural language processing (NLP) (such as the Transformer architecture) to time series prediction may not fully consider the unique time patterns and periodic characteristics inherent in retail business operations (such as weekly periodicity, seasonal effects, etc.). Many time series prediction models modified based on the Transformer architecture have been proposed, such as Informer, PatchTST, AutoFormer. These models cannot efficiently and fully utilize the rich features in the retail people, goods, and store dimensions, nor do they handle well the relationship between the variable to be predicted (sales volume) and other features. This neglect of domain characteristics may lead to insufficient feature representation and loss of key periodic information, thereby limiting the accuracy of prediction. The TimeXer model proposes an efficient learning method that combines endogenous variables and exogenous variables, but the original model only supports numerical features and does not have a module for processing retail heterogeneous features.
[0004] Therefore, there is a need for improved methods and systems for predicting retail sales volume. Summary of the Invention
[0005] The embodiments of this specification are made to solve the above existing technical problems, and their purpose is to improve the prediction of retail sales volume.
[0006] According to a first aspect of the present disclosure, there is provided a computer-implemented method for predicting retail sales volume, including: obtaining multi-source heterogeneous feature data related to a retail business, where the multi-source heterogeneous feature data at least includes sequence data representing historical endogenous variables and feature data representing exogenous variables; performing feature encoding processing on the multi-source heterogeneous feature data to generate a vectorized representation segment of a set of endogenous variables based on the sequence data representing historical endogenous variables and an encoded vector of a set of exogenous variables based on the feature data representing exogenous variables; initializing at least one learnable global token vector; inputting the vectorized representation segment of the endogenous variables, the encoded vector of the exogenous variables, and at least one learnable global token vector into a core prediction model, where the core prediction model passes through one or more attention mechanism layers, enabling the global token vector to have attention interactions with the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables respectively, and updating the representation of the global token vector through the attention interaction to integrate information from the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables; predicting the future sales volume of the retail business based on the final representation of the global token vector obtained after being processed by the core prediction model.
[0007] According to a further embodiment of the present disclosure, performing feature encoding processing on the multi-source heterogeneous feature data further includes: performing numerical embedding processing on numerical features included in the multi-source heterogeneous feature data to obtain corresponding encoded vectors or a part of the vectorized representation segment; wherein, the numerical embedding processing includes using a learnable linear transformation matrix to expand the numerical features from their original dimension to a predetermined vector dimension.
[0008] According to a further embodiment of the present disclosure, performing feature encoding processing on the multi-source heterogeneous feature data further includes: performing categorical embedding processing on categorical features included in the multi-source heterogeneous feature data to obtain corresponding encoded vectors or a part of the vectorized representation segment; wherein, the categorical embedding processing includes: based on the categorical features, looking up corresponding basic embedding vectors from a preset embedding dictionary; applying a gating function to process the basic embedding vectors.
[0009] According to a further embodiment of the present disclosure, the gating function is a learnable function or a predetermined function calculated based on the categorical features themselves.
[0010] According to a further embodiment of the present disclosure, when generating the vectorized representation segment and / or the encoded vector, perform at least one of the following operations: performing peer feature aggregation on embedding features from the same original feature group that are used to form the vectorized representation segment or the encoded vector; performing cross-level feature splicing on embedding features from different original feature groups or types that are used to form the vectorized representation segment or the encoded vector.
[0011] According to a further embodiment of the present disclosure, the feature encoding process further includes: using at least one sequence encoding technique configured to adapt to the periodic characteristics of the retail business to process the temporal part in the sequence data representing historical endogenous variables and / or the feature data of exogenous variables, so as to generate corresponding vectorized representation segments representing endogenous variables and / or encoded vectors of exogenous variables.
[0012] According to a further embodiment of the present disclosure, the sequence encoding technique includes applying at least one convolutional layer, and the convolutional kernel size, sliding step, and / or convolutional kernel moving dimension of the convolutional layer are configured to adapt to the periodic characteristics of the retail business.
[0013] According to a further embodiment of the present disclosure, the sequence encoding technique further includes at least one of the following: applying at least one recurrent neural network layer; applying at least one convolutional neural network layer; applying at least one multi-layer perceptron layer; fusing the outputs from different feature extraction paths to form a final encoded vector or vectorized representation segment.
[0014] According to a further embodiment of the present disclosure, one or more attention mechanism layers at least partially implement attention interaction through a multi-head hidden attention mechanism.
[0015] According to a further embodiment of the present disclosure, the core prediction model includes at least one processing block, and at least one processing block is configured to: perform the following operations to update the representation of the global token vector: perform a first attention interaction between the global token vector and the vectorized representation segment of the endogenous variable; for the representation of at least the global token vector updated by the first attention interaction, perform a second attention interaction between it and the encoded vector of the exogenous variable; after the first attention interaction and / or the second attention interaction, process the representation of at least the global token vector through a mixture-of-experts network.
[0016] According to a further embodiment of the present disclosure, the vectorized representation segment of the endogenous variable is generated by dividing the sequence data representing historical endogenous variables into multiple fixed-length segments and vectorizing each segment.
[0017] According to a further embodiment of the present disclosure, it further includes: when it is necessary to predict the more long-term future sales volume beyond the single prediction range of the core prediction model, a recursive prediction method is adopted, where the recursive prediction method includes combining the future sales volume prediction result obtained in the previous prediction stage with the corresponding future known exogenous variable feature data as part of the input historical data for the next prediction stage, and iteratively performing the prediction.
[0018] According to a further embodiment of the present disclosure, the multi-source heterogeneous feature data includes at least one or a combination of more than one of the following: historical sales volume sequence, product identifier, store identifier, promotion information, environmental characteristics, member feature statistics of purchased products, time-related features, and product inventory information.
[0019] According to a second aspect of the present disclosure, there is provided a method for training a core prediction model to predict retail sales volume, including: obtaining multi-source heterogeneous feature data related to retail business and corresponding sequence data representing historical endogenous variables as training data, where the multi-source heterogeneous feature data at least includes a part of the original historical endogenous variable sequence data for generating a vectorized representation segment of the endogenous variable, and the original exogenous variable feature data for generating an encoded vector of the exogenous variable; using the training data to train the core prediction model, where the training process includes: performing feature encoding processing on the multi-source heterogeneous feature data in the training data to generate a vectorized representation segment of the endogenous variable and an encoded vector of the exogenous variable; providing the vectorized representation segment of the endogenous variable, the encoded vector of the exogenous variable, and at least one initialized learnable global token vector to the core prediction model; through one or more attention mechanism layers of the core prediction model, enabling the global token vector to perform attention interaction with the vectorized representation segment of the endogenous variable and the encoded vector of the exogenous variable respectively, and updating the representation of the global token vector through the attention interaction to integrate information from the vectorized representation segment of the endogenous variable and the encoded vector of the exogenous variable, so as to obtain an updated representation of the global token vector for learning; based on the comparison between the prediction of the sequence data representing historical endogenous variables by the core prediction model using the updated representation of the global token vector and the actual sequence data representing historical endogenous variables, iteratively adjusting the learnable parameters of the core prediction model so that the trained core prediction model can predict the future sales volume of the retail business based on the final representation of the global token vector.
[0020] According to a further embodiment of the present disclosure, the step of obtaining training data includes: constructing a plurality of training samples from the sequence data representing historical endogenous variables and the multi-source heterogeneous feature data in a sliding window manner, each training sample including a historical data part of a first predetermined length and a corresponding target sequence part of a second predetermined length, where the first predetermined length and the second predetermined length depend on the periodic characteristics of the retail business.
[0021] According to a further embodiment of the present disclosure, the historical data part of each training sample further includes future known exogenous variable feature data corresponding to the second predetermined length.
[0022] According to a further embodiment of the present disclosure, it further includes: applying progressive random masking to at least a part of the original historical endogenous variable sequence data part, where the proportion of random masking gradually decreases with the training process; applying a causal masking mechanism.
[0023] According to a third aspect of the present disclosure, there is provided a system for predicting retail sales volume, including: at least one processor; at least one memory coupled to the at least one processor, on which computer-executable instructions are stored, and when the computer-executable instructions are executed by the at least one processor, the system executes the method provided by the present disclosure.
[0024] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer-executable instructions are stored, and when the computer-executable instructions are executed, the method provided by the present disclosure is executed.
[0025] Compared with the prior art, the technical solution of the present disclosure can achieve one or more of the following technical effects:
[0026] Through a specially designed feature encoding processing flow, the present disclosure can effectively process and transform multi-source heterogeneous feature data from different sources with different types and structures. By introducing a learnable global token vector inside the core prediction model and enabling it to perform deep, multi-stage attention interactions with the vectorized representation segments representing historical endogenous variables and the encoded vectors representing exogenous variables, effective integration and condensation of these multi-source heterogeneous information are achieved. In addition, through numerical embedding processing of numerical features and categorical embedding processing of categorical features, the original heterogeneous input can be transformed into a unified vector representation.
[0027] The present disclosure realizes the adaptability to the periodic characteristics of the retail business, and can better capture and utilize the complex spatio-temporal dependence relationships and long-term periodic patterns hidden in the data. This explicit consideration and adaptation of business periodicity enable the model to more accurately grasp the fluctuation laws of sales volume over time (such as weekly, monthly, quarterly).
[0028] Due to the more effective processing and fusion of heterogeneous data, better adaptation to business periodicity above, and the global token-guided information integration mechanism focused on the prediction target, the prediction model adopting the technical solution of the present disclosure can achieve higher accuracy in actual retail sales volume prediction tasks. Compared with some baseline models, the method of the present disclosure can significantly reduce key prediction error metrics, such as the mean absolute percentage error (MAPE).
[0029] The core prediction model of the present disclosure can integrate components such as the multi-head hidden attention mechanism (MLA) and the mixture of experts network (MoE), which not only improves the accuracy of the solution of the present disclosure, but also has the practicability, scalability and economy for deployment and application in actual large-scale retail scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The illustrative aspects of the present application are described in detail below with reference to the following drawings:
[0031] Figure 1 is a schematic diagram of the core process of a method for predicting retail sales volume according to an embodiment of the present disclosure;
[0032] Figure 2 is a more detailed schematic diagram of a feature encoding processing unit according to an example of the present disclosure;
[0033] Figure 3 is a schematic diagram of the internal architecture of a core prediction model according to an example of the present disclosure;
[0034] Figure 4 is a conceptual structural schematic diagram of a multi-head hidden attention mechanism (MLA) module according to an example of the present disclosure;
[0035] Figure 5 is a conceptual structural schematic diagram of a mixture of experts network (MoE) module according to an example of the present disclosure;
[0036] Figure 6 is a schematic diagram of the core process of a method for training a core prediction model according to an example of the present disclosure;
[0037] Figure 7 is a schematic diagram of a sliding window sample construction process according to an example of the present disclosure;
[0038] Figure 8 is a schematic diagram of a process adopting a recursive prediction method according to an example of the present disclosure;
[0039] Figure 9 is a hardware structure block diagram of a computing device that can be used to implement the method of the present disclosure according to an example of the present disclosure;
[0040] Figure 10 is a high-level functional module block diagram of an overall prediction system for predicting retail sales volume according to an example of the present disclosure;
[0041] Figure 11 is a high-level functional module block diagram of an overall training system for training a core prediction model according to an example of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without making creative efforts fall within the scope of protection of the present disclosure.
[0043] In addition, in the description of the present disclosure, unless otherwise clearly specified and limited, terms such as "arranged", "installed", "connected", "coupled", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.
[0044] In addition, the term "and / or" involved in the embodiments of the present disclosure is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, terms such as "first", "second", etc. (if any) in this specification, the claims, and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here.
[0045] As described in the background art section, in the field of retail sales volume prediction, there is still a need for improvement in the prior art in effectively processing and integrating data from diverse sources and heterogeneous types, as well as in capturing and utilizing the time patterns (such as periodicity) unique to the retail business. Failure to fully address these challenges may lead to limited performance of the prediction model, affecting the accuracy and reliability of the prediction results.
[0046] To address one or more of the above technical challenges, the present disclosure provides an improved method and system for predicting retail sales volume. By performing specially designed feature encoding processing on the obtained multi-source heterogeneous feature data and using a core prediction model, a learnable global token vector is deeply interacted with the encoded endogenous variable information and exogenous variable information through a specific attention mechanism, and prediction is made based on the final representation of the global token vector, thereby being able to more effectively integrate information and adapt to the complex characteristics of the retail business, and further improving the accuracy of the prediction.
[0047] The method, system, and related aspects for predicting retail sales according to exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0048] Referring to Figure 1 , which shows a schematic flowchart of a method 100 for predicting retail sales according to an embodiment of the present disclosure. The method 100 may be implemented by one or more computing devices (e.g., Figure 9 the computing device 900 shown).
[0049] The method 100 includes: operation 102 of obtaining multi-source heterogeneous feature data 10 related to the retail business. The multi-source heterogeneous feature data 10 at least includes sequence data representing historical endogenous variables and feature data representing exogenous variables.
[0050] The "retail business" herein refers to the sales activities of goods or services to end consumers, such as, but not limited to, supermarkets, department stores, specialty stores, or online retail platforms, etc.
[0051] The multi-source heterogeneous feature data 10 refers to data collected from different sources and having different types and structures. The sources of the multi-source heterogeneous feature data 10 may include, for example, enterprise internal systems such as point-of-sale (POS) systems, inventory management systems, customer relationship management (CRM) systems, and promotion management systems; or may also include external data sources such as weather service APIs, public holiday data, macroeconomic indicator data, etc.
[0052] In an exemplary configuration, the multi-source heterogeneous feature data may include at least one of: product identifier, store identifier, historical sales volume sequence, promotion information, and environmental features; the environmental features may include at least one of temperature, wind force, and weather condition. Further, the multi-source heterogeneous feature data 10 at least includes numerical features and categorical features. For example, the numerical features may include historical sales volume values, prices, and temperatures; the categorical features may include product SKU codes, store IDs, and weather condition codes.
[0053] Preferably, the multi-source heterogeneous feature data 10 is classified according to a predefined classification system. For example, the multi-source heterogeneous feature data 10 can be divided into: a basic feature layer, a time-series feature layer, and an environmental feature layer. The basic feature layer mainly includes static features or semi-static features. For example, it can include features such as product SKU codes and store IDs. The basic feature layer can include numerical features and categorical features. The time-series feature layer mainly includes dynamic features. For example, it can include dynamic data recorded with timestamps, such as historical sales volume sequences, promotion intensity time series, etc. The time-series feature layer can include numerical features and categorical features. The environmental feature layer can include static features, semi-static features, or dynamic features. The environmental feature layer can include: numerical environmental features, such as the temperature gradient calculated based on the daily maximum temperature and minimum temperature (e.g., ΔT = daily maximum temperature - daily minimum temperature) or wind force level (e.g., based on the Beaufort scale); categorical environmental features, such as weather states encoded through a predefined dictionary (e.g., the dictionary can contain 10 or more meteorological categories, such as "heavy rain" encoded as "01", "dust storm" encoded as "02", etc.); and composite environmental features, such as the holiday impact index calculated by combining the pedestrian flow coefficient and the economic location weight (e.g., index H = log(pedestrian flow coefficient) × economic location weight).
[0054] The multi-source heterogeneous feature data 10 can be obtained from one or more of the aforementioned data sources. For example, database queries, API calls, or file reading can be used. Preferably, the obtained data should cover a sufficiently long historical period (e.g., 420 days or more of history) to support model training and contain future information relevant to the prediction period (such as known promotion plans, weather forecasts, etc.).
[0055] In the multi-source heterogeneous feature data 10, the sequence data representing historical endogenous variables refers to the data whose historical values directly constitute the target time series to be predicted, and whose future values are the outputs of model predictions. In the retail sales prediction scenario of the present disclosure, the most typical sequence data of historical endogenous variables is the time series of historical sales volume, such as the daily, weekly, or monthly sales quantity or sales amount of a specific product at a specific store or channel.
[0056] The feature data representing exogenous variables refers to other relevant feature information that can affect the changes of historical endogenous variables (such as future sales volume in this example), but whose own changes are usually not directly determined by or mainly determined by the historical values of historical endogenous variables. These exogenous variables provide auxiliary context, driving factors, or constraints for the prediction model to understand and predict the changes of historical endogenous variables. Examples of the feature data of exogenous variables can include the aforementioned promotion information, environmental features, static product attributes, static store attributes, holiday information, member feature statistics, time-related features, and product inventory information, etc.
[0057] Method 100 further includes: operation 104, performing feature encoding processing on the multi-source heterogeneous feature data 10 to generate a vectorized representation segment 110 of a set of endogenous variables based on the sequence data representing historical endogenous variables and an encoded vector 120 of a set of exogenous variables based on the feature data representing exogenous variables.
[0058] Feature encoding processing in the present disclosure refers to a series of processes that convert the original, diverse multi-source heterogeneous feature data 10 into a structured numerical representation suitable for input into a subsequent core prediction model (such as Figure 3 the core prediction model 300 shown in). The goal of this processing is to extract meaningful features and represent them as vectors or sequences of vectors in a unified format.
[0059] A vectorized representation segment 110 of a set of endogenous variables is a plurality of vectors formed by performing specific processing on the sequence data representing historical endogenous variables. These vector segments are designed to capture local and temporal information of the historical endogenous variable sequence. For example, as will be described in detail later with reference to Figure 7 A sequence containing historical sales data for a specific number of days (e.g., 28 days) can be divided into multiple fixed-length segments (Patches), and each segment is then encoded into a vector of a fixed dimension, thereby forming a vectorized representation segment 110 of a set of endogenous variables containing multiple vectors.
[0060] An encoded vector 120 of a set of exogenous variables is a plurality of vectors formed by performing encoding processing on the feature data representing exogenous variables. Each encoded vector can correspond to one or a set of exogenous variable features. For example, different exogenous variable features such as weather conditions, promotion activity information, holiday information, etc., after being processed by embedding and / or specific sequence encoding techniques, can be respectively represented as one or more encoded vectors of a fixed dimension, jointly constituting an encoded vector 120 of a set of exogenous variables. These encoded vectors carry external context information that affects the change of the target variable.
[0061] In an exemplary implementation, the process of performing the feature encoding processing to generate a vectorized representation segment 110 of endogenous variables and an encoded vector 120 of exogenous variables may further include one or more of the following sub-processes. For more detailed exemplary implementations of these sub-processes, reference may be made to the description of the feature encoding processing unit 200 and its internal components (such as the feature embedding subunit 210 and the sequence encoding and preliminary feature construction subunit 220) in the following (detailed schematic diagram of feature encoding processing): Figure 2 (Description of the feature encoding processing unit 200 and its internal components, such as the feature embedding subunit 210 and the sequence encoding and preliminary feature construction subunit 220, shown in the detailed schematic diagram of feature encoding processing):
[0062] Numerical embedding is performed on the numerical features contained in the multi-source heterogeneous feature data 10 to obtain corresponding encoding vectors or to form a part of the vectorized representation fragment. The numerical embedding process may include using a learnable linear transformation matrix to expand the numerical features from their original dimensions to a predetermined vector dimension.
[0063] A category embedding process is performed on the category features contained in the multi-source heterogeneous feature data 10 to obtain a corresponding encoding vector or to form a part of a vectorized representation fragment. The category embedding process may include: searching for a corresponding basic embedding vector from a preset embedding dictionary based on the category features, and applying a gating function to process the basic embedding vector. The gating function may be a learnable function or a predetermined function calculated based on the category features themselves.
[0064] At least one sequence encoding technique configured to adapt to the cyclical characteristics of the retail business is used to process the sequence data representing the historical endogenous variables and / or the time series part of the feature data representing the exogenous variables to generate the corresponding vectorized representation fragment 110 of the endogenous variables and / or the encoding vector 120 of the exogenous variables. The sequence encoding technique may include applying at least one convolution layer (whose convolution kernel size, sliding step size and / or convolution kernel shift dimension are configured to adapt to the cyclical characteristics of the retail business), or applying at least one recurrent neural network layer, or applying at least one convolution neural network layer, or applying at least one multi-layer perceptron layer, or fusing the outputs from different feature extraction paths.
[0065] The method 100 further includes: operation 106 , initializing at least one learnable global token vector 130 .
[0066] The learnable global token vector 130 in the present disclosure refers to a vector whose parameters are learned and optimized during the model training process. Unlike vectors representing specific input features (such as specific historical sales values or specific weather conditions), the global token vector 130 usually does not carry specific input information when initialized, but serves as a "floating", plastic representation carrier. For example, it can be randomly initialized or use some preset initial value.
[0067] Initialization is to set an initial state or value for the at least one learnable global token vector 130. Its dimension is usually consistent with the vector dimension (e.g., dimension d) in the vectorized representation fragment 110 of the endogenous variable and the encoding vector 120 of the exogenous variable generated by the aforementioned feature encoding process, so as to facilitate effective attention interaction and information integration in the core prediction model later.
[0068] The role of the at least one learnable global token vector 130 will be in the subsequent core prediction model processing (such as Figure 1 Operation 108 in , andFigure 3 It is specifically described in the core prediction model 300 shown. It will be designed to capture and compress the global context information, the overall sequence pattern, and the complex relationships between different feature sources related to predicting future sales volume by performing deep attention interactions with the vectorized representation segment 110 of the endogenous variable and the encoded vector 120 of the exogenous variable. Its "learnable" property means that the model will adjust the parameters of this global token vector during training so that it can be used for better information integration and final prediction.
[0069] Method 100 further includes: operation 108, inputting the vectorized representation segment 110 of the endogenous variable, the encoded vector 120 of the exogenous variable, and at least one learnable global token vector 130 into the core prediction model (such as the core prediction model 300 described below). The core prediction model enables the global token vector 130 to perform attention interactions with the vectorized representation segment 110 of the endogenous variable and the encoded vector 120 of the exogenous variable respectively through one or more attention mechanism layers, and updates the representation of the global token vector 130 through the attention interaction to integrate the information from the vectorized representation segment 110 of the endogenous variable and the encoded vector 120 of the exogenous variable. The ultimate goal of this operation is to obtain the final representation 150 of the global token vector that is information-compressed and can be used directly for sales volume prediction.
[0070] The core prediction model in this disclosure refers to a computational model designed to deeply process multiple provided input vectors (i.e., the vectorized representation segment 110 of the endogenous variable, the encoded vector 120 of the exogenous variable, and the global token vector 130), and learn the relationships between them, extract key patterns, and ultimately drive the prediction through an internal complex interaction mechanism. An exemplary detailed architecture of this core prediction model can be referred to the description of the Figure 3 core prediction model 300 later.
[0071] The core prediction model includes one or more attention mechanism layers. These attention mechanism layers are configured to perform specific types of attention interactions.
[0072] One type of attention interaction enables the global token vector 130 to perform attention interaction with the vectorized representation segment 110 of the endogenous variable. This interaction can be, for example, self-attention interaction, where the global token and the endogenous segment jointly serve as part of the query, key, and value. This attention interaction allows the global token to extract information from the historical patterns, trends, and periodicity of the endogenous variable, and at the same time enables the representation of the endogenous segment to take into account the global context.
[0073] Another type of attention interaction enables the global token vector 130 (e.g., the updated representation after the above-mentioned interaction) to have an attention interaction with the encoded vector 120 of the exogenous variable. This interaction can be, for example, a cross-attention interaction, where the global token (e.g., its updated representation) can serve as the query, and the encoded vector of the exogenous variable serves as the key and value. This attention interaction allows the global token to further integrate information from external influencing factors such as weather, promotions, holidays, etc.
[0074] Through the attention interaction, the representation of the global token vector 130 is gradually and iteratively updated to integrate the information from the vectorized representation segments 110 of the endogenous variable and the encoded vector 120 of the exogenous variable. This means that during the processing of the core prediction model, the global token vector 130 is gradually updated from its initial state to a condensed representation that can reflect multi-faceted information related to future sales prediction.
[0075] In a preferred example, the core prediction model can, for example, consist of multiple modified Transformer Encoder Blocks. As Figure 3 shown, the core prediction model 300 can include at least one processing block (e.g., Figure 3 the encoder block 310_i in it), which is configured to perform the above-mentioned attention interaction. For example, the processing block can be configured to: perform a first attention interaction (e.g., self-attention interaction) between the global token vector and the vectorized representation segments of the endogenous variable; and for the representation of at least the global token vector updated by the first attention interaction, perform a second attention interaction (e.g., cross-attention interaction) between it and the encoded vector of the exogenous variable. This phased interaction helps the model to orderly fuse information from different sources.
[0076] Preferably, one or more attention mechanism layers can at least partially implement the attention interaction through the multi-head latent attention mechanism (MLA). MLA helps to efficiently process sequence information and capture complex dependencies by introducing a set of learnable latent vectors as information transfer and bottlenecks. For a more detailed exemplary structure of MLA, reference can be made to the subsequent description of Figure 4 this.
[0077] Preferably, inside one or more processing blocks of the core prediction model, after the first attention interaction and / or the second attention interaction, the representation of at least the global token vector can also be processed by a mixture of experts network (MoE). MoE helps to improve the capacity and efficiency of the model by dynamically selecting a part of the expert sub-networks to process the input. For a more detailed exemplary structure of MoE, reference can be made to the subsequent description of Figure 5 this.
[0078] Method 100 further includes: operation 112, based on the final representation 150 of the global token vector obtained after processing by the core prediction model, predicting the future sales volume 40 of the retail business.
[0079] The final representation 150 of the global token vector is the final state reached by the global token vector 130 inside the core prediction model after passing through one or more processing layers (e.g., the L repeated encoder blocks 310_i shown in Figure 3 ), as well as multiple attention interactions and possible mixture-of-experts network processing that occur therein. The final representation 150 is designed to have integrated the key information related to predicting future sales volume from the vectorized representation segments 110 of the endogenous variables and the encoded vectors 120 of the exogenous variables.
[0080] After obtaining the final representation 150 of the global token vector, this representation is used to generate a specific predicted value for the future sales volume of the retail business. This is typically achieved by inputting the final representation 150 into one or more output layers. For example, as shown in Figure 3 , the end of the core prediction model 300 (i.e., the core prediction model) may include a projection layer 330. The projection layer 330 receives the final representation 150 of the global token vector (whose dimension is, for example, d) and linearly or non-linearly maps it to the target prediction dimension. For example, if it is necessary to predict the daily sales volume for the next 7 days, the projection layer 330 can map the d-dimensional global token final representation to a 7-dimensional vector, where each element of the vector corresponds to a sales volume prediction value for a future day, thereby obtaining the final future sales volume 40.
[0081] The output future sales volume 40 can be point prediction values for one or more future time steps. In one example, it can be the daily sales volume prediction for consecutive days in the future. In another example, these per-time-step prediction values can also be further aggregated as needed to provide prediction results at different time granularities, such as the total sales volume for the next week or the average daily sales volume for the next month, etc.
[0082] In some preferred examples, if it is necessary to predict the future sales volume over a longer term that exceeds the single prediction range of the core prediction model (e.g., the future sales volume over a longer period that cannot be covered by forward propagation), a recursive prediction method can be further adopted. As will be described in detail later with reference to Figure 8 , this recursive prediction method includes combining the future sales volume prediction result obtained in the previous prediction stage (which at least partially serves as the historical endogenous variables in the next stage) with the corresponding future known exogenous variable feature data as part of the input historical data for the next prediction stage, and iteratively performing the prediction operation. In this way, the time range of the prediction can be gradually extended, thereby allowing for a longer-term future sales volume.
[0083] Refer toFigure 2 , which shows a more detailed schematic diagram of the feature encoding processing unit 200 according to an example of the present disclosure. The feature encoding processing unit 200 is designed to implement the function of operation 104 in the foregoing method 100, that is, based on the acquired multi-source heterogeneous feature data 10, generate a vectorized representation segment 110 of a set of endogenous variables and an encoded vector 120 of a set of exogenous variables for use by the core prediction model (such as Figure 3 the core prediction model 300 shown). The feature encoding processing unit 200 can be implemented, for example, in the feature encoding module 1020 of a prediction system (such as Figure 10 the prediction system 1000 shown).
[0084] As Figure 2 shown, the feature encoding processing unit 200 can receive the multi-source heterogeneous feature data 10 as its initial input. In one example, the feature encoding processing unit 200 can include at least two main sub-processing stages or sub-units: a feature embedding sub-unit 210, and a sequence encoding and preliminary feature construction sub-unit 220.
[0085] The feature embedding sub-unit 210 is used to perform a preliminary vectorization conversion on the original features in the input multi-source heterogeneous feature data 10. In a specific implementation, the feature embedding sub-unit 210 can perform the following operations:
[0086] For the numerical features included in the multi-source heterogeneous feature data 10 (such as Figure 2 the numerical feature 2122 schematically input to the numerical processing channel 212 in
[0087] v_num = ReLU(W·x + b)
[0088] Where v_num is the output numerical embedding feature 2126, and its dimension is d.
[0089] Optionally, before performing the linear transformation, other preprocessing operations can also be performed on the numerical features, such as normalization processing, etc.
[0090] For the categorical features included in the multi-source heterogeneous feature data 10 (such as Figure 2 the categorical feature 2142 schematically input to the categorical processing channel 214), the categorical embedding processing can be performed on it through the categorical processing channel 214 to obtain the corresponding categorical embedding feature (such as the categorical embedding feature 2142). The categorical feature 2142 is, for example, a discrete category, identifier, or attribute (such as, product SKU code, store ID, promotion activity type code, specific code for weather conditions, etc.).
[0091] The categorical processing channel 214 may include an embedding dictionary lookup unit 2144. The embedding dictionary lookup unit 2144 is associated with one or more embedding dictionaries E. In an exemplary implementation, this categorical embedding processing includes: based on the categorical feature 2142 (such as its index index(x)), looking up the corresponding basic embedding vector (such as E[index(x)]) from the preset embedding dictionary E. The embedding dictionary E is usually a learnable matrix, and its dimension can be designed as, for example, R^{C×d}, where C represents the maximum category cardinality that the categorical feature may have (for example, in a specific implementation, C can be set to be greater than or equal to 100 to support categorical features with a large number of unique values), and d is the target embedding dimension.
[0092] The categorical processing channel 214 may further include a gating function application unit 2146 for dynamically adjusting or weighting the basic embedding vector obtained from the embedding dictionary lookup unit 2144. The gating function application unit 2146 can calculate the gating signal (or weight vector), denoted as g(x), based on the original categorical feature 2142 itself (or information derived from it). The gating function can be a learnable function or a predetermined function calculated based on the categorical feature itself.
[0093] The categorical processing channel 214 may further include a combination unit 2148, which is used to interact the gating signal g(x) with the basic embedding vector E[index(x)], for example, through element-wise multiplication (denoted as ). Therefore, the final categorical embedding feature v_cat can be calculated as follows:
[0094] v_cat = E[index(x)] g(x)
[0095] This gating mechanism enables the model to dynamically modulate its embedding representation according to the specific values or context of the input categorical features.
[0096] The numerical embedding features and categorical embedding features output by the feature embedding subunit 210 (schematically represented as embedding features 222 in Figure 2 ) are then passed to the sequence encoding and preliminary feature construction subunit 220.
[0097] The sequence encoding and preliminary feature construction subunit 220 is used to further process the embedded features (embedding features 222) to finally generate a vectorized representation segment 110 of a set of endogenous variables input to the core prediction model and a set of encoded vectors 120 of exogenous variables. The processing of this sequence encoding and preliminary feature construction subunit 220 may include, for example:
[0098] Fragmentation (patching) and encoding of the endogenous variable sequence (e.g., implemented by the internal endogenous sequence fragmentation and encoding unit 224): For the sequence data representing historical endogenous variables (after embedding), it can be divided into multiple fixed-length segments, and each segment is further vectorized and encoded to form a vectorized representation segment 110 of a set of endogenous variables. For example, the embedding of a historical sales volume sequence for a specific number of days can be divided into multiple segments of shorter time windows, and each segment is encoded as a vector of a fixed dimension. For example, see Figure 3 , an endogenous variable sequence of every 28 days can be used as a segment, including a set of 7 endogenous vectors, each vector including 4 days of data (such as sales data), which can be mapped to a d-dimensional vector through an MLP (as described below). A more specific example of this process can be referred to the description of Figure 7 later.
[0099] Sequence / feature encoding of exogenous variables (e.g., implemented by the internal exogenous sequence / feature encoding unit 226): For the feature data representing exogenous variables (especially the time series part thereof, and also including the embedding of static features), specific encoding techniques can be used to convert it into a set of encoded vectors 120 of exogenous variables.
[0100] During the above encoding process, at least one sequence encoding technique configured to adapt to the periodic characteristics of the retail business can be adopted to process the time series part in the sequence data of historical endogenous variables and / or the feature data of exogenous variables. The sequence encoding technique may include applying at least one convolutional layer (such as convolutional layer 2262, whose convolutional kernel size, stride, and / or convolutional kernel movement dimension are configured to adapt to the periodic characteristics of the retail business), or applying at least one recurrent neural network layer (such as recurrent neural network layer 2264), or applying at least one multi-layer perceptron (MLP) layer (such as multi-layer perceptron layer 2266). In some examples, alternatively or additionally, at least one convolutional neural network layer (not shown in the figure) can also be applied.
[0101] In some examples, if multiple different feature extraction paths are adopted for the same feature group or different feature groups, the outputs from different paths can be fused by the optional feature extraction path fusion unit 2268 to form a more robust or information-rich encoded vector or vectorized representation segment.
[0102] Furthermore, inside the sequence encoding and preliminary feature construction subunit 220, optional preliminary aggregation operations can also be performed: for example, an aggregation operation is performed on the vectors from the same logical feature group by the optional peer feature aggregation unit 232; and / or the feature vectors from different feature layers or types are concatenated by the optional cross-level feature concatenation unit 234. These operations help to construct components of a better-structured intermediate feature representation before feeding it into the core model.
[0103] For example, peer feature aggregation (which can be achieved by the peer feature aggregation unit 232) can aggregate the embedding features from the same predefined level or logical group (for example, all weather-related embedding features). In one example, specifically, for a set of weather-related embedding features, such as temperature embedding v_temp, wind force embedding v_wind, and weather condition embedding v_condition, they can be aggregated into a comprehensive weather feature vector v_weather by summation:
[0104] v_weather = Σ(v_temp + v_wind + v_condition)
[0105] Similarly, other embedding features belonging to the same logical group can also perform similar peer aggregation operations.
[0106] Cross-level feature concatenation (which can be achieved through the cross-level feature concatenation unit 234) can combine the embedding representations from different predefined feature layers (e.g., the base feature layer, the temporal feature layer, and the environmental feature layer). For example, the embedding features representing the static commodity base attributes (assumed to be represented as v_static, which can be formed through peer aggregation), the embedding features representing the dynamic historical sales time series (represented as v_dynamic, which can also be formed through peer aggregation), and the embedding features representing environmental factors (such as the aforementioned v_weather or other environmental feature embeddings v_env, which can also be formed through peer aggregation) can be concatenated along their feature dimensions. This operation combines these feature vectors from different sources and levels into a more comprehensive multi-level feature representation, i.e., Concat(v_static, v_dynamic, v_env)).
[0107] Finally, the sequence encoding and the preliminary feature construction subunit 220 output a set of vectorized representation segments 110 of endogenous variables and a set of encoded vectors 120 of exogenous variables, which will be input (along with the initialized global token) into Figure 3 the core prediction model 300 shown.
[0108] Refer to Figure 4 , which shows a conceptual structural schematic diagram of the MLA module 400 according to an example of the present disclosure. As described above, the attention mechanism layers (such as the first attention mechanism 314_i and / or the second attention mechanism 318_i) inside the core prediction model (such as Figure 3 the core prediction model 300 shown) can at least partially implement their attention interaction functions through this MLA module 400. The MLA module 400 introduces a fixed number of learnable "latent vectors" (assumed to be the number L). The attention calculation of the MLA module 400 no longer directly occurs between the input sequence tokens, but is relayed through the latent vectors. The L latent vectors act as a learnable and fixed-size information bottleneck or summary. Regardless of the length of the input sequence (how large N is), the number L of these latent vectors is fixed. The model learns how to extract and compress key information from the large input into these L latent vectors. The self-attention between the latent vectors (L x L) allows these latent vectors to interact with each other and refine their internal representations within a fixed dimension, forming an independent core processing module independent of the input length.
[0109] By using the MLA module, the complexity of processing long sequences is reduced to O(N*L) or O(L²), approaching linear level, while providing an efficient fixed-size information bottleneck to process information from large-scale inputs, significantly improving the scalability of the model and enabling it to be applied to process data with ultra-high dimensions or ultra-long sequences.
[0110] For details about the MLA module 400, please refer to the following papers published by DeepSeek-AI et al.: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (arXiv:2405.04434) and DeepSeek-V3 Technical Report (arXiv:2412.19437).
[0111] According to experiments, by using MLA to replace MHA in the traditional Transformer model, the prediction accuracy for all categories (1 - MAPE) can be increased by approximately 0.5%.
[0112] Refer to Figure 5 , which shows a conceptual structural diagram of the MoE module 500 (e.g., it can be the DeepSeek MoE module) according to an example of the present disclosure. As described above, within each processing block (e.g., the encoder block 310_i) inside the core prediction model (e.g., Figure 3 the core prediction model 300 shown), the MoE module 500 can be used to replace the traditional Feed-Forward Network (FFN) part (as shown by the DeepSeek MoE module 312_i in Figure 3 ). The MoE module 500 aims to achieve a higher model capacity and performance at a lower computational cost by dynamically selecting a subset of expert sub-networks for processing the input, and is particularly suitable for processing and transforming the representation of key information such as global token vectors.
[0113] As Figure 5 shown, the MoE module 500 can divide experts into shared experts and task-specific experts. Shared experts are responsible for extracting general features across tasks; while task-specific experts are optimized for specific domains or tasks.
[0114] The expert screening and routing part adopts a finer-grained division of experts, groups the experts and selects the activated experts through a two-stage routing mechanism: first, the group of experts with the highest scores is screened out through a gating network, and then the top-K experts are selected within this group.
[0115] For some details about the MoE module 500, reference can also be made to the following papers published by DeepSeek-AI et al.: DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (arXiv:2405.04434) and DeepSeek-V3 Technical Report (arXiv:2412.19437).
[0116] By adopting this MoE module 500 that includes shared experts, task-specific experts, and a two-stage Top-K routing mechanism, the core prediction model 300 of the present disclosure can more effectively utilize a large number of parameters (achieving high capacity through numerous experts) while maintaining a low single-sample inference calculation cost (because only a few experts are activated each time). This structure helps the model learn more complex and fine-grained feature representations, thereby improving prediction performance. For example, in some application scenarios, introducing such an MoE module can bring an increase of approximately 1.2% in the prediction accuracy of all product categories (measured by 1 - MAPE).
[0117] Refer to Figure 6 , which shows a schematic flowchart of a method 600 for training a core prediction model (such as the core prediction model 300 shown in Figure 3 ) to predict retail sales according to an example of the present disclosure. The optimized core prediction model 300 obtained by training through the method 600 can be used to implement the aforementioned prediction method 100. The method 600 aims to utilize the patterns and regularities in historical data to adjust and optimize the internal learnable parameters of the core prediction model 300.
[0118] As shown in Figure 6 , the method 600 includes: operation 602, obtaining multi-source heterogeneous feature data 10 related to the retail business and the corresponding sequence data representing historical endogenous variables as training data. The multi-source heterogeneous feature data 10 at least includes a part of the original historical endogenous variable sequence data for generating a vectorized representation segment of the endogenous variable, and the original exogenous variable feature data for generating an encoded vector of the exogenous variable. As described above for operation 102 in the method 100, the multi-source heterogeneous feature data 10 can contain various types and sources of information, and the sequence data representing historical endogenous variables (such as historical sales) provides the target value that the model needs to learn to predict in a supervised learning framework. In operation 602, in order to prepare the training data, for example, Figure 7The sliding window sample construction process 700 shown constructs a training sample with a fixed length from the original time series data. In addition, preprocessing steps may be performed on the acquired original data, such as but not limited to missing value filling (such as using triple exponential smoothing) and feature standardization (such as using RobustScaler and other methods).
[0119] The method 600 further includes: operation 604, using the training data to train the core prediction model 300, wherein the training process includes:
[0120] Sub-operation 6042: Perform feature encoding processing on the multi-source heterogeneous feature data 10 in the training data to generate a vectorized representation segment 110 of the endogenous variable and an encoding vector 120 of the exogenous variable. The feature encoding processing can refer to the above description of Figure 2 The feature encoding processing unit 200 shown is implemented by performing numerical embedding processing and category embedding processing (wherein the category embedding processing may include applying a gating function), and using a sequence encoding technology (such as convolution, recurrent neural network, multi-layer perceptron or a fusion thereof) adapted to the cycle characteristics of the retail business to process the time series part, and may include same-level feature aggregation and / or cross-level feature concatenation to form an input for the core model.
[0121] Sub-operation 6044: providing the vectorized representation fragment 110 of the endogenous variable, the encoding vector 120 of the exogenous variable, and at least one initialized learnable global token vector 130 to the core prediction model 300 .
[0122] Sub-operation 6046: Through one or more attention mechanism layers of the core prediction model 300, the global token vector 130 is made to interact with the vectorized representation fragment 110 of the endogenous variable and the encoding vector 120 of the exogenous variable respectively, and the representation of the global token vector is updated through the attention interaction to integrate the information from the vectorized representation fragment 110 of the endogenous variable and the encoding vector 120 of the exogenous variable, thereby obtaining an updated representation of the global token vector for learning. The internal processing of the core prediction model 300 can refer to the above description of Figure 3 Preferably, the attention mechanism layer can be at least partially implemented by a multi-head latent attention mechanism (MLA) (e.g. Figure 4 ), and its processing block can be implemented through a mixture of experts network (MoE) after attention interaction (as shown in Figure 5 ) processes a representation of at least the global token vector.
[0123] The updated representation of the global token vector for learning (and the corresponding historical endogenous variable sequence data as actual values) is then used in the core iteration loop of the training process, such as Figure 6The part shown from operations 608 to 614. In a typical training iteration:
[0124] At operation 608, a specific training strategy can be applied. In a preferred example, especially when the core prediction model 300 is based on a modified Transformer decoder block, this operation further includes: applying progressive random masking to at least a part of the original historical endogenous variable sequence data part, where the proportion of random masking gradually decreases with the training process (for example, the masking rate ρ can linearly decay from about 0.5 to about 0.1); and applying a causal masking mechanism (for example, by applying a causal masking matrix M to ensure that when calculating the attention weights at each time step, information from subsequent time steps is not attended to).
[0125] At operation 610, the prediction of the sequence data of historical endogenous variables can be made based on the core prediction model 300 using the updated representation of the global token vector (for example, the predicted 7-day future sales volume, as shown in Figure 3 340), and compared with the actual sequence data of historical endogenous variables (the actual 7-day future sales volume, as shown in Figure 3 350) to calculate the loss value. The loss function can be a function suitable for regression tasks, such as mean squared error (MSE) or mean absolute error (MAE).
[0126] At operation 612, based on the comparison between the prediction of the sequence data representing historical endogenous variables made using the updated representation of the global token vector by the core prediction model and the actual sequence data representing historical endogenous variables, the learnable parameters of the core prediction model are iteratively adjusted so that the trained core prediction model can predict the future sales volume of the retail business based on the final representation of the global token vector. This parameter adjustment process typically uses the backpropagation algorithm and an optimizer (such as Adam, SGD, etc., and the optimizer can be configured with an exemplary initial learning rate, such as about 1e-4) to update the model parameters.
[0127] This parameter adjustment process is repeatedly performed in the training iteration, and at operation 614, it is checked whether the training process meets a preset termination condition (for example, reaching the maximum number of training epochs, or meeting an early stopping condition such as no significant improvement in performance for consecutive epochs). If the termination condition is not met, the next iteration continues; if it is met, the training process ends.
[0128] By repeatedly executing the above steps, the parameters of the core prediction model 300 are optimized.
[0129] Preferably, a sliding window approach is adopted to construct multiple training samples from the sequence data representing historical endogenous variables and multi-source heterogeneous feature data. Each training sample includes a historical data part of a first predetermined length and a corresponding target sequence part of a second predetermined length, where the first predetermined length and the second predetermined length depend on the periodic characteristics of the retail business. Refer to Figure 7 , which shows a schematic diagram of a sliding window sample construction process 700 according to an example of the present disclosure. This process 700 is an effective method for constructing training samples or validation / test samples from raw time series data (e.g., the time series part of the multi-source heterogeneous feature data 10 obtained in operation 602 of method 600 and the corresponding historical sales data), and is particularly suitable for time series prediction tasks.
[0130] As Figure 7 shown, the sliding window sample construction process 700 generally receives a relatively long segment of raw time series data 702 as input. This data may include, for example, historical feature values arranged in chronological order and corresponding target values (such as historical sales).
[0131] This process defines a sliding window 704 with a fixed total length. In Figure 7 a preferred example, the total length of this sliding window 704 can be set to 35 days. The sliding window 704 moves along the raw time series data 702 according to a preset step size (e.g., sliding day by day, i.e., the step size is 1 day, as shown by step size 710).
[0132] For each position of the sliding window 704, the data within the window is divided into at least two parts:
[0133] One part is a historical data part 706 of a first predetermined length (also referred to as the X part or input feature part). For example, in a 35-day window, the data for the first 28 days (i.e., the first predetermined length) (including the observations of various relevant multi-source heterogeneous feature data 10 within these 28 days) can be used as the historical data part 706.
[0134] The other part is a corresponding target sequence part 708 of a second predetermined length (also referred to as the Y part or label part). For example, in the above 35-day window, the sequence data (such as sales) of the actual historical endogenous variables for the 7 days (i.e., the second predetermined length) immediately following the historical data part 706 can be used as the target sequence part 708.
[0135] In some other examples, other sliding window settings can be adopted. Preferably, the first predetermined length and the second predetermined length can depend on the periodic characteristics of the retail business, such as weekly cycle, monthly cycle, quarterly cycle, etc. For example, the total length of the sliding window can be set to 8 days, where the data of the first 7 days is used as the historical feature part, and the actual sales data of the last 1 day is used as the target sequence part. Another example is that the total length of the sliding window can be set to 4 months, where the data of the first 3 months is used as the historical feature part, and the actual sales data of the last 1 month is used as the target sequence part.
[0136] In some more preferred examples, the historical data part 706 (i.e., input features) of each training sample can also include future known exogenous variable feature data corresponding to the second predetermined length (target sequence period). These future known exogenous variable feature data refer to external influencing factors whose values are known in advance during the time period when the target sequence part 708 occurs, such as weather forecasts, scheduled promotion activity information, or determined holiday arrangements during this time period. Taking these future known information as part of the input features helps the model to more accurately predict the target sequence. As shown in the prediction process 712, using 28 days of historical data part and 7 days of future known exogenous variable feature data, future sales can be predicted online.
[0137] By moving the sliding window 704 along the entire original time series data 702, a large number of training / validation samples that may have partial overlaps can be generated. The samples include historical information (and possibly future known information) as input features, and the corresponding future target sequence as the prediction label. These generated samples can then be used for training (e.g., in method 600) or evaluating the time series prediction model (e.g., Figure 3 the core prediction model 300 shown).
[0138] This way of constructing samples with a sliding window helps the model learn the dynamic patterns and time dependencies in the time series data.
[0139] Referring to Figure 8 , which shows a recursive prediction process 800 for predicting longer-term future sales beyond the single prediction range of the core prediction model using a recursive prediction method according to an example of the present disclosure. This recursive prediction method can be applied when performing the foregoing method 100 when a sales prediction with a longer time span (e.g., future 14 days, 21 days or longer) than the prediction that can be generated by one forward propagation of the core prediction model 300 (e.g., future 7 days) is required.
[0140] As Figure 8 shown, the process 800 generally involves iteratively using the core prediction model 300 (in Figure 8It may be represented as multiple calls to the core prediction model, such as 300_1, 300_2, etc.
[0141] In the first prediction stage of recursive prediction:
[0142] First, prepare the initial input data 802. The initial input data 802 includes a sequence of data representing historical endogenous variables within a specific historical time span (e.g., the most recent 28 days, denoted as [D - 28, D]) and related multi-source heterogeneous feature data, as well as future known exogenous variable feature data corresponding to the first short-term prediction period (e.g., the next 7 days, denoted as [D, D + 7]).
[0143] Then, provide the initial input data 802 to the core prediction model (the first call, i.e., core prediction model call 300_1).
[0144] Core prediction model call 300_1 processes this input and outputs the future sales volume prediction result 804 (e.g., y) for the first short-term prediction period [D, D + 7].
[0145] In the subsequent prediction stages of recursive prediction (e.g., the second prediction stage for predicting the sales volume in [D + 7, D + 14]):
[0146] First, construct the partial input historical data 806 for the next prediction stage. This typically includes: combining the future sales volume prediction result 804 (i.e., y) obtained in the previous prediction stage (i.e., the first prediction stage) with the corresponding future known exogenous variable feature data (i.e., the known exogenous features corresponding to the time period [D, D + 7], if these features are processed together with the sales volume in the model, or simply treating the predicted sales volume y as the historical endogenous variable for the next stage).
[0147] Then, splice or integrate this combined data (representing the information that has occurred or been predicted during the period [D, D + 7]) with earlier and still relevant actual historical data (e.g., the actual historical data for the time period [D - 21, D]) to form a new historical data window that conforms to the input requirements of the core prediction model 300 (e.g., still 28 days in length, but shifted 7 days into the future in time, covering [D - 21, D + 7]). At the same time, it is also necessary to prepare the future known exogenous variable feature data corresponding to the new short-term prediction period (e.g., [D + 7, D + 14]).
[0148] Provide the newly formed input data to the core prediction model (the second call, i.e., core prediction model call 300_2).
[0149] The core prediction model invokes 300_2 to process this new input and outputs the future sales volume prediction result 808 (e.g., y') for the second short-term prediction period [D+7, D+14].
[0150] By iteratively performing the steps of the subsequent prediction phase described above (i.e., feeding back the latest predicted sales volume result as part of the input for the next round), the predicted time range can be gradually extended to the desired longer term. The multiple short-term prediction results finally obtained can be concatenated to form a prediction of the entire long-term future sales volume.
[0151] This recursive prediction method enables a core model capable of accurate short-term prediction to be effectively applied to long-term prediction tasks, expanding the applicability of the model.
[0152] Refer to Figure 9 , which shows a block diagram of an exemplary computing device 900 that can be used to implement the methods (e.g., prediction method 100 or training method 600) or systems (e.g., prediction system 1000 or training system 1100) according to some embodiments of the present disclosure. The computing device 900 can be various forms of computing devices, such as, but not limited to, a personal computer (PC), a server, a workstation, a laptop computer, a tablet computer, or a dedicated data processing device or an embedded system. The computing device 900 provides a suitable hardware platform for executing the methods described in the present disclosure and implementing the described systems.
[0153] As Figure 9 shown, the computing device 900 generally includes at least one processor 910. The processor 910 can be a central processing unit (CPU), a microprocessor (MPU), a graphics processing unit (GPU) (particularly suitable for performing parallel computing-intensive machine learning tasks), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a neural network processing unit (NPU), a tensor processing unit (TPU), or any other control logic circuit capable of processing data and executing instructions or a combination thereof. The processor 910 is responsible for executing computer-executable instructions stored in the memory to implement various functions and operations described in the present disclosure.
[0154] The computing device 900 also includes at least one memory 920, which is communicatively coupled to the processor 910 (e.g., via one or more system buses not explicitly shown in Figure 9 ). The memory 920 can be any type of volatile memory (e.g., random access memory RAM, such as DRAM, SRAM) for temporarily storing the instructions being executed by the processor 910 and the data being processed while the computing device 900 is running.
[0155] The computing device 900 generally also includes a persistent storage device 930. The persistent storage device 930 is used for long-term storage of data and program instructions, and can retain information even after the device is powered off or restarted. The persistent storage device 930 can be, for example, a hard disk drive (HDD), a solid state drive (SSD), an optical disc drive, a flash memory device, or other types of non-volatile computer-readable storage media. The operating system, application software (e.g., a software program containing one or more modules implementing the prediction method 100 or the training method 600 described in this disclosure), and various data required when executing these methods (e.g., input multi-source heterogeneous feature data 10, intermediate data generated during model training such as vectorized representation segments 110 of endogenous variables, encoded vectors 120 of exogenous variables, final representations 150 of global token vectors, trained core prediction model parameters, etc.) can be stored in the persistent storage device 930 and loaded into the main memory 920 for access and execution by the processor 910 when needed.
[0156] The computing device 900 may also include one or more input / output (I / O) interfaces 940. The input / output interface 940 provides a way for the computing device 900 to interact with an operator or other external devices or systems. For example, through the input / output interface 940, the computing device 900 can be connected to:
[0157] One or more network interfaces 950 (e.g., Ethernet cards, Wi-Fi modules) for communicating with other computing devices, database servers, or external data sources through a network (such as the Internet, local area network), for example, for obtaining multi-source heterogeneous feature data 10 or transmitting the final future sales volume 40 prediction results.
[0158] One or more display devices 960 (e.g., monitors, touchscreens) for presenting information to a user or operator, such as prediction results, monitoring information of the training process, etc.
[0159] One or more input devices 970 (e.g., keyboards, mice, touchpads, microphones) for receiving instructions, configuration parameters, or other input information from a user or operator.
[0160] In addition, the computing device 900 may also include other components well-known to those skilled in the art, such as a power supply unit, various internal buses (such as address buses, data buses, control buses) for connecting the above-mentioned main components, and other dedicated hardware accelerators, etc. (not all shown in Figure 9 for the sake of brevity).
[0161] In the context of the present disclosure, the memory 920 and / or the persistent storage device 930 (as one or more computer-readable storage media) can store computer-executable instructions. When these computer-executable instructions are executed by the processor 910, the computing device 900 can be caused to implement all or part of the operations of the aforementioned prediction method 100, or implement all or part of the operations of the training method 600. Similarly, each functional module of the prediction system (such as Figure 10 the prediction system 1000 shown) or the training system (such as Figure 11 the training system 1100 shown) can be implemented by executing corresponding software instructions on the computing device 900.
[0162] It should be understood that Figure 9 the architecture of the computing device 900 shown is only exemplary and is not intended to be limiting. Those skilled in the art can modify, add to, or expand the configuration of the computing device according to specific application requirements and technological developments. For example, a system including multiple processor cores or multiple physical processors can be adopted, a distributed storage architecture can be used, a dedicated AI coprocessor or FPGA can be integrated for hardware acceleration, etc. As long as these variations can implement the core technical solutions proposed by the present disclosure, they should all be considered to fall within the protection scope of the present disclosure.
[0163] Referring to Figure 10 , which shows a high-level functional module block diagram of a prediction system 1000 for predicting retail sales according to an embodiment of the present disclosure. The prediction system 1000 is configured to implement the prediction method 100 described above with reference to Figure 1 . The prediction system 1000 can be implemented on one or more computing devices (such as Figure 9 the computing device 900 shown), and its various functional modules can be constructed by hardware, software, or a combination thereof.
[0164] As Figure 10 shown, in an exemplary configuration, the prediction system 1000 can include the following main functional modules:
[0165] Data acquisition and preprocessing module 1010: used to obtain multi-source heterogeneous feature data 10 required for performing the prediction task from various internal or external data sources. As described above for operation 102 in method 100, these data can include sequence data of historical endogenous variables and feature data representing exogenous variables, and these data at least include numerical features and categorical features. In addition, the data acquisition and preprocessing module 1010 can also perform necessary data preprocessing operations, such as data cleaning, format conversion, missing value processing, and feature standardization or normalization, etc., to prepare data 10 suitable for subsequent processing. In some implementations, if the preparation of training data and prediction data involves sliding window construction (such as with reference toFigure 7 For the described process 700), related functions can also be undertaken or called by a part of the data acquisition and preprocessing module 1010. The core function of the data acquisition and preprocessing module 1010 corresponds to the data acquisition and preliminary preparation part of operation 102 in method 100.
[0166] Feature encoding module 1020: It is used to receive the multi-source heterogeneous feature data 10 (which may have been preliminarily processed) from the data acquisition and preprocessing module 1010, and is responsible for performing the feature encoding process in operation 104 of method 100. Its purpose is to convert the input multi-source heterogeneous feature data 10 into a structured input for the core prediction model, that is, a vectorized representation segment 110 of a set of endogenous variables and an encoded vector 120 of a set of exogenous variables. In a specific implementation, the feature encoding module 1020 may internally contain a functional unit for implementing feature embedding (the detailed structure and processing can refer to the description of the feature embedding subunit 210 in the feature encoding processing unit 200 shown above, including numerical embedding and categorical embedding) and a functional unit for implementing sequence encoding and preliminary feature construction (the detailed structure and processing can refer to the description of the sequence encoding and preliminary feature construction subunit 220 in the feature encoding processing unit 200 shown above, including using sequence encoding techniques adapted to the characteristics of the retail business cycle for the time series part, such as convolution, recurrent neural network, or multi-layer perceptron, and possible peer feature aggregation and / or cross-level feature splicing). The output of the feature encoding module 1020, that is, the vectorized representation segment 110 of the endogenous variables and the encoded vector 120 of the exogenous variables, is passed to the core prediction model module 1030. Figure 2 For the description of the feature embedding subunit 210 in the feature encoding processing unit 200 shown above, including numerical embedding and categorical embedding), and a functional unit for implementing sequence encoding and preliminary feature construction (the detailed structure and processing can refer to the description of the sequence encoding and preliminary feature construction subunit 220 in the feature encoding processing unit 200 shown above, including using sequence encoding techniques adapted to the characteristics of the retail business cycle for the time series part, such as convolution, recurrent neural network, or multi-layer perceptron, and possible peer feature aggregation and / or cross-level feature splicing). Figure 2 The output of the feature encoding module 1020, that is, the vectorized representation segment 110 of the endogenous variables and the encoded vector 120 of the exogenous variables, is passed to the core prediction model module 1030.
[0167] Core prediction model module 1030: It corresponds to or implements the core prediction model in method 100 (its more detailed internal architecture is as shown in the core prediction model 300). The core prediction model module 1030 receives the vectorized representation segment 110 of the endogenous variables and the encoded vector 120 of the exogenous variables from the feature encoding module 1020, and one or more learnable global token vectors 130 initialized at this stage or inside it. The core prediction model module 1030 is responsible for performing the core prediction processing logic defined in operation 108 of method 100, that is, through one or more attention mechanism layers inside it (for example, it can be at least partially implemented through a multi-head latent attention mechanism (MLA) (as shown above)), so that the global token vectors interact with the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables respectively, and update the representation of the global token vectors through these interactions to integrate information from all parties. In some implementations, the processing blocks inside this module (for example, as shown above) Figure 3 For the core prediction model 300 shown above). The core prediction model module 1030 receives the vectorized representation segment 110 of the endogenous variables and the encoded vector 120 of the exogenous variables from the feature encoding module 1020, and one or more learnable global token vectors 130 initialized at this stage or inside it. Figure 4 For the multi-head latent attention mechanism (MLA) (as shown above)), so that the global token vectors interact with the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables respectively, and update the representation of the global token vectors through these interactions to integrate information from all parties. Figure 3The encoder block 310_i) shown can also further process the representation of the global token vector through a mixture of experts network (MoE) (as shown in Figure 5 ). The final output of the core prediction model module 1030 is the final representation 150 of the global token vector after in-depth processing and information condensation.
[0168] Prediction result output module 1040: It is used to receive the final representation 150 of the global token vector from the core prediction model module 1030 and is responsible for performing the final prediction step defined in operation 112 of method 100. That is, based on the final representation 150 of this global token vector, the future sales volume 40 of the retail business is predicted through an appropriate output layer (such as the projection layer 330 in Figure 3 ). The prediction result output module 1040 can also be responsible for presenting or transmitting the predicted future sales volume 40 (schematically marked as prediction output in Figure 10 ) in a format suitable for users or downstream applications (such as inventory management, business intelligence reports, etc.).
[0169] It should be emphasized that Figure 10 the modular structure shown is only an exemplary functional division, aiming to clearly illustrate how the prediction system 1000 implements each main stage of the prediction method 100. In a specific engineering implementation, the functional boundaries of these modules may be different, the functions of some modules may be merged into other modules, or further divided into more sub-modules. For example, some or all of the functions of the feature encoding module 1020 and the core prediction model module 1030 are usually constructed as an end-to-end neural network model in a deep learning-based implementation.
[0170] Referring to Figure 11 , it shows a high-level functional module block diagram of a training system 1100 for training a core prediction model (such as the core prediction model 300 shown in Figure 3 ) to predict retail sales volume. The training system 1100 is configured to implement the training method 600 described above with reference to Figure 6 . Similar to the prediction system 1000, the training system 1100 can also be implemented on one or more computing devices (such as the computing device 900 shown in Figure 9 ), and its various functional modules can be constructed by hardware, software, or a combination thereof.
[0171] As shown in Figure 11 , in an exemplary configuration, the training system 1100 may include the following main functional modules:
[0172] Training data acquisition and preprocessing module 1110: It is used to obtain and preprocess data from various internal or external data sources ( Figure 11The training data acquisition and preprocessing module 1110 (schematically labeled as the training data source in FIG.) collects the training data required for model training. The function of this module 1110 corresponds to operation 602 in the training method 600, that is, obtaining multi-source heterogeneous feature data 10 related to the retail business (including numerical features and categorical features) and the corresponding sequence data representing historical endogenous variables (as the training target). The training data acquisition and preprocessing module 1110 may also perform necessary data preprocessing operations, such as data cleaning, missing value filling (e.g., using triple exponential smoothing), feature standardization (e.g., using RobustScaler), and constructing training samples with a fixed length from the original time series data through a sliding window technique (an exemplary implementation of which can refer to Figure 7 the sliding window sample construction process 700 shown). This module 1110 provides the prepared training data (schematically represented in Figure 11 as the training data containing data 10) to the model training engine 1120.
[0173] Model training engine 1120: Used to execute the main training logic and iterative optimization process in the training method 600. The model training engine 1120 is responsible for using the training data to train the core prediction model 300. Its training process utilizes the optimized feature representations obtained by performing specific processing on the multi-source heterogeneous feature data 10 in the training data, and iteratively adjusts the learnable parameters of the core prediction model 300 based on the comparison between the prediction and the actual historical data (corresponding to the core training steps in the training method 600, including feature encoding, core model processing, loss calculation, and parameter adjustment, etc.). In an exemplary implementation, the model training engine 1120 may internally include a feature encoding processing unit and a core model training unit.
[0174] The feature encoding processing unit corresponds to or implements the feature encoding module 1020 in the prediction system 1000 (or Figure 2 the feature encoding processing unit 200 shown), and is used to perform feature embedding, sequence encoding (adaptable to the characteristics of the retail business cycle), and possibly preliminary aggregation on the multi-source heterogeneous feature data 10 in the input training data to generate a vectorized representation segment 110 of a set of endogenous variables and a set of encoded vectors 120 of exogenous variables for the core model to learn.
[0175] The core model training unit is used to manage the training iteration of the core prediction model 300 (the trainable version). It provides the vectorized representation segment 110 of the endogenous variables, the encoded vectors 120 of the exogenous variables, and at least one initialized learnable global token vector 130 generated by the feature encoding processing unit to the core prediction model 300. It is also responsible for coordinating the attention interaction and the update process of the global token representation inside the core prediction model 300 (such as referring to Figure 3Description). During training, this unit can also apply specific training strategies, such as applying progressive random masking and causal masking mechanisms during the decoder processing stage of the core prediction model 300 (if based on the Transformer architecture). This unit is also responsible for calculating the loss between the model prediction and the actual historical target value, and updating the parameters of the core prediction model 300 based on this loss through an optimization algorithm.
[0176] Model parameter storage 1130: Used to store all learnable parameters of the core prediction model 300 during training, such as the weight matrix and bias vector of the neural network. During training, the model training engine 1120 (especially its core model training unit) continuously interacts with the model parameter storage 1130, reads the current parameters for forward propagation, and writes the new parameters back after parameter update. After training is completed, the finally optimized model parameters are also saved in this storage 1130 for subsequent model evaluation or deployment.
[0177] Optional post-training model output module 1140: It can be called after the successful end of the training process to export or package the fully trained and optimized core prediction model 300 (or its parameters) stored in the model parameter storage 1130 in an appropriate format, so as to facilitate its deployment to an actual prediction application (for example, deployed to Figure 10 the core prediction model module 1030 of the prediction system 1000 as shown), or for further analysis and evaluation.
[0178] Similar to the prediction system 1000, Figure 11 the module division of the training system 1100 as shown is also an exemplary functional organization method, and its specific implementation can be adjusted according to application requirements.
[0179] To further improve the performance and efficiency of the prediction method (such as method 100) and prediction system (such as prediction system 1000) proposed in this disclosure during actual deployment and application, some optimization measures in the prediction inference stage can be considered. These optimization measures are not an inherent part of the core prediction algorithm or model architecture, but help to improve the practicality and responsiveness of the entire prediction solution.
[0180] For example, when performing predictive inference, Mixed-Precision Inference can be implemented. This means that in some computational layers of the core prediction model (such as 300), specifically those that are computationally intensive but do not require extremely high numerical precision, like certain large-scale matrix operations, a lower-precision floating-point representation (such as FP16 half-precision) can be used, while in other critical parts that are sensitive to precision (such as layers that require high-precision accumulation or the final output layer), a higher precision (such as FP32 single-precision) is maintained. By combining different precisions such as FP16 and FP32 for calculations, it is possible to effectively reduce the memory footprint during model inference while ensuring that the prediction accuracy is not significantly affected, and it can also leverage the acceleration capabilities of modern computing hardware (such as GPUs supporting tensor cores) for low-precision operations, thereby improving the prediction throughput and reducing the latency of a single prediction.
[0181] To accelerate the response speed in real-time or high-frequency prediction scenarios, a Feature Caching Mechanism can be deployed. For intermediate features or raw features that are relatively time-consuming to compute during the feature encoding process (such as those executed by the feature encoding processing unit 200 or the feature encoding module 1020), but whose computed results may remain unchanged or change slowly within a short period (such as certain complex derived features, or semi-static features retrieved from external data sources with a low update frequency), they can be stored in the cache after the first calculation or retrieval. When subsequent prediction requests arrive, if the relevant original input has not changed or the change is within an acceptable range, the system can quickly retrieve these pre-computed features from the cache without having to repeat the time-consuming calculation or query operations. This feature caching mechanism can significantly reduce the time required for feature preparation, thereby improving the efficiency of the overall prediction process and the real-time response ability of the system.
[0182] In addition, to provide decision-makers with more comprehensive information to evaluate the reliability of prediction results and assist in risk management, the prediction system 1000 (such as in its prediction result output module 1040) can construct or integrate an Uncertainty Quantification Module. This module can provide a corresponding confidence interval or probability distribution for the point prediction value of the future sales volume 40 generated based on the output of the core prediction model 300, or by using other statistical methods (such as Monte Carlo dropout techniques, constructing Bayesian neural network models, etc.). Outputting the confidence interval of the prediction (such as the 95% confidence interval) helps users understand the potential fluctuation range and uncertainty level of the prediction results, enabling them to better consider potential risks and make more robust decisions when making key business decisions such as inventory planning, marketing budget allocation, and supply chain adjustment.
[0183] These optimization measures in the prediction and inference phase can be selectively configured and implemented according to specific application scenarios, performance requirements, and available computing resources to further enhance the overall effect and value of the technical solution of the present disclosure in practical applications.
[0184] Compared with the prior art, the method, system, and related aspects for predicting retail sales volume provided by the present disclosure can bring at least one of the following significant technical advantages and beneficial effects.
[0185] First, by adopting specially designed feature encoding processing (including fine-grained feature embedding processing, such as gated embedding mechanisms for categorical features and normalization and linear transformation for numerical features) and subsequent feature aggregation processing (for example, introducing learnable global tokens inside the core prediction model and enabling them to perform deep attention interactions with different types of encoded features, as well as supporting peer aggregation and cross-level splicing of features), the technical solution of the present disclosure can more effectively process and fuse multi-source heterogeneous feature data. This enables the model to extract richer and more valuable signals from complex and diverse input information, overcoming the common limitations of traditional methods in data fusion.
[0186] Second, the present disclosure emphasizes and realizes the adaptability to the periodic characteristics of the retail business. This is reflected at multiple levels: for example, in the feature encoding stage, convolutional layers configured with specific convolutional kernels (whose size, stride, or moving dimension adapt to the business cycle) or other sequence encoding techniques can be used to process time-series features; inside the core prediction model, through advanced attention mechanisms such as the multi-head latent attention mechanism (MLA), the model can capture and utilize the complex spatio-temporal dependencies and long-term periodic patterns hidden in the data. This explicit consideration and adaptation to business periodicity enable the model to more accurately grasp the fluctuation patterns of sales volume over time (such as weekly, monthly, or seasonally).
[0187] Due to the more effective fusion of heterogeneous data and better adaptation to business periodicity as described above, the prediction model adopting the technical solution of the present disclosure can achieve higher prediction accuracy in actual retail sales volume prediction tasks. As observed in some exemplary validations, compared with some baseline models, the method of the present disclosure can significantly reduce key prediction error metrics, such as the mean absolute percentage error (MAPE). In a specific validation scenario, the MAPE can be reduced to approximately 30%, which may represent a relative performance improvement of approximately 20% compared with traditional methods or deep learning models that have not been specifically optimized. The improvement in prediction accuracy has direct economic benefits and operational value for retail enterprises in aspects such as inventory management, reducing out-of-stock and overstock situations, optimizing promotion strategies, and enhancing customer satisfaction.
[0188] In addition, the technical solution of the present disclosure, especially the design of its core prediction model (e.g., based on the Transformer architecture and modified to include components such as MLA and MoE), considers the scalability and computational efficiency of the model while improving performance. For example, the MLA mechanism helps handle long-sequence inputs, while the MoE mechanism can reduce the computational cost of a single inference while maintaining the high capacity of the model. These features enable the solution of the present disclosure to not only achieve breakthroughs in accuracy but also possess the practicality and robustness for deployment and application in actual large-scale retail scenarios. Its modular design (e.g., feature encoding, core model, optional recursive prediction, inference optimization, etc.) also provides flexibility for adjustment and expansion according to different business requirements.
[0189] Through innovations at multiple levels such as feature representation, model architecture, and training / inference strategies, the technical solution of the present disclosure effectively addresses the deficiencies of existing retail sales prediction technologies in handling complex data and adapting to business characteristics, thereby being able to provide more accurate, reliable, and practically valuable prediction results, providing strong technical support for the refined operation and intelligent decision-making of retail enterprises.
[0190] Those skilled in the art will readily think of other advantages and modifications. Therefore, in a broader sense, the present invention is not limited to the specific details and representative embodiments shown and described herein. Thus, modifications can be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents.
Claims
1. A computer-implemented method for predicting retail sales volume, characterized in that, Including: Obtaining multi-source heterogeneous feature data related to the retail business, where the multi-source heterogeneous feature data at least includes sequence data representing historical endogenous variables and feature data representing exogenous variables; Performing feature encoding processing on the multi-source heterogeneous feature data to generate a vectorized representation segment of a set of endogenous variables based on the sequence data representing historical endogenous variables and an encoded vector of a set of exogenous variables based on the feature data representing exogenous variables; Initializing at least one learnable global token vector; Inputting the vectorized representation segment of the endogenous variables, the encoded vector of the exogenous variables, and the at least one learnable global token vector into a core prediction model. The core prediction model passes through one or more attention mechanism layers, enabling the global token vector to have attention interactions with the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables respectively, and updating the representation of the global token vector through the attention interaction to integrate information from the vectorized representation segment of the endogenous variables and the encoded vector of the exogenous variables; And Predicting the future sales volume of the retail business based on the final representation of the global token vector obtained after being processed by the core prediction model.
2. The method according to claim 1, wherein The performing feature encoding processing on the multi-source heterogeneous feature data further includes: Performing numerical embedding processing on the numerical features included in the multi-source heterogeneous feature data to obtain corresponding encoded vectors or a part of the vectorized representation segment; Wherein, the numerical embedding processing includes using a learnable linear transformation matrix to expand the numerical feature from its original dimension to a pre-determined vector dimension.
3. The method according to claim 1, wherein The performing feature encoding processing on the multi-source heterogeneous feature data further includes: Performing categorical embedding processing on the categorical features included in the multi-source heterogeneous feature data to obtain corresponding encoded vectors or a part of the vectorized representation segment; Wherein, the categorical embedding processing includes: Based on the categorical feature, looking up the corresponding basic embedding vector from a preset embedding dictionary; and Applying a gating function to process the basic embedding vector.
4. The method according to claim 3, wherein The gating function is a learnable function or a predetermined function calculated based on the categorical feature itself.
5. The method according to claim 1, wherein When generating the vectorized representation segment and / or the encoded vector, perform at least one of the following operations: Performing peer feature aggregation on the embedding features from the same original feature group that are used to form the vectorized representation segment or the encoded vector; and Performing cross-level feature splicing on the embedding features from different original feature groups or types that are used to form the vectorized representation segment or the encoded vector.
6. The method according to claim 1, characterized in that, The feature encoding processing further includes: Using at least one sequence encoding technique configured to adapt to the periodic characteristics of the retail business to process the time series part in the sequence data representing historical endogenous variables and / or the feature data representing exogenous variables, so as to generate the corresponding vectorized representation segment of the endogenous variables and / or the encoded vector of the exogenous variables.
7. The method according to claim 6, wherein The sequence encoding technique includes applying at least one convolutional layer, and the convolutional kernel size, stride, and / or convolutional kernel movement dimension of the convolutional layer are configured to adapt to the periodic characteristics of the retail business.
8. The method according to claim 6, wherein The sequence encoding technique further includes at least one of the following: Applying at least one recurrent neural network layer; Applying at least one convolutional neural network layer; Applying at least one multi-layer perceptron layer; and Fusing the outputs from different feature extraction paths to form a final encoded vector or the vectorized representation segment.
9. The method according to claim 1, wherein The one or more attention mechanism layers at least partially implement the attention interaction through a multi-head latent attention mechanism.
10. The method according to claim 1, wherein The core prediction model includes at least one processing block, and the at least one processing block is configured to: Perform the following operations to update the representation of the global token vector: perform a first attention interaction between the global token vector and the vectorized representation segment of the endogenous variable; and for at least the representation of the global token vector updated by the first attention interaction, perform a second attention interaction between it and the encoded vector of the exogenous variable; And After the first attention interaction and / or the second attention interaction, process at least the representation of the global token vector through a mixture-of-experts network.
11. The method according to claim 1, characterized in that, The vectorized representation segment of the endogenous variable is generated by dividing the sequence data representing the historical endogenous variables into multiple fixed-length segments and vectorizing each segment.
12. The method according to claim 1, wherein Further includes: When predicting the longer-term future sales volume beyond the single prediction range of the core prediction model is required, a recursive prediction method is adopted, where the recursive prediction method includes combining the future sales volume prediction result obtained in the previous prediction stage with the corresponding future known exogenous variable feature data as part of the input historical data for the next prediction stage, and iteratively performing the prediction.
13. The method according to claim 1, wherein The multi-source heterogeneous feature data includes at least one or more combinations of the following: historical sales volume sequence, product identifier, store identifier, promotion information, environmental features, member feature statistics of purchased products, time-related features, and product inventory information.
14. A method for training a core prediction model to predict retail sales volume, characterized in that, Includes: Obtaining multi-source heterogeneous feature data related to the retail business and the corresponding sequence data representing the historical endogenous variables as training data, where the multi-source heterogeneous feature data at least includes the original historical endogenous variable sequence data part for generating the vectorized representation segment of the endogenous variable and the original exogenous variable feature data for generating the encoded vector of the exogenous variable; Using the training data to train the core prediction model, where the training includes: Performing feature encoding processing on the multi-source heterogeneous feature data in the training data to generate the vectorized representation segment of the endogenous variable and the encoded vector of the exogenous variable; Providing the vectorized representation segment of the endogenous variable, the encoded vector of the exogenous variable, and at least one initialized learnable global token vector to the core prediction model; Through one or more attention mechanism layers of the core prediction model, the global token vector respectively performs attention interaction with the vectorized representation segments of the endogenous variables and the encoded vectors of the exogenous variables, and updates the representation of the global token vector through the attention interaction to integrate the information from the vectorized representation segments of the endogenous variables and the encoded vectors of the exogenous variables, so as to obtain the updated representation of the global token vector for learning; And Based on the comparison between the prediction of the sequence data representing the historical endogenous variables by the core prediction model using the updated representation of the global token vector and the actual sequence data representing the historical endogenous variables, by iteratively adjusting the learnable parameters of the core prediction model, so that the trained core prediction model can predict the future sales volume of the retail business based on the final representation of the global token vector.
15. The method according to claim 14, wherein Obtaining multi-source heterogeneous feature data related to the retail business and the corresponding sequence data representing the historical endogenous variables as training data includes: Constructing a plurality of training samples from the sequence data representing the historical endogenous variables and the multi-source heterogeneous feature data in a sliding window manner, each training sample including a historical data part of a first predetermined length and a corresponding target sequence part of a second predetermined length, wherein the first predetermined length and the second predetermined length depend on the periodic characteristics of the retail business.
16. The method according to claim 15, wherein The historical data part of each training sample further includes future known exogenous variable feature data corresponding to the second predetermined length.
17. The method according to claim 14, wherein Further includes: Applying progressive random masking to at least a part of the original historical endogenous variable sequence data part, wherein the proportion of random masking gradually decreases with the training process; And Applying a causal masking mechanism.
18. A system for predicting retail sales volume, characterized in that, Includes: At least one processor; And At least one memory coupled to the at least one processor, the memory stores computer-executable instructions, and when the computer-executable instructions are executed by the at least one processor, the system executes the method according to any one of claims 1-13.
19. A computer-readable storage medium, on which computer-executable instructions are stored, and when the computer-executable instructions are executed, the method according to any one of claims 1-13 is executed.