Training method and device of sales volume prediction model, electronic equipment and storage medium
By constructing multiple scenario-based expert modules and target mapping relationships to train the sales forecasting model, the problem of insufficient stability and adaptability of traditional sales forecasting methods in complex scenarios is solved, and more refined and accurate product sales forecasting is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-03-17
AI Technical Summary
Traditional sales forecasting methods are insufficient to fully depict the changing patterns under different business scenarios, resulting in insufficient stability and adaptability of forecast results in complex business scenarios, making it difficult to meet the complex forecasting needs in actual business.
By constructing multiple pre-defined expert modules, a target mapping relationship is built based on the feature information of sales data samples. The initial model is then trained using the target loss function to generate a sales prediction model, including scenario-based expert modules for global, major promotions, new products, and old products, enabling adaptive learning for different scenarios.
It improves the robustness and scenario adaptability of the sales forecasting model, enabling it to accurately capture sales trends and cyclical patterns in different scenarios, thereby improving forecast accuracy and generalization ability.
Smart Images

Figure CN121685013A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of model training technology, and in particular to a training method, apparatus, electronic device and storage medium for a sales forecasting model. Background Technology
[0002] In product sales scenarios, sales volume is influenced by a variety of factors, resulting in a complex, diverse, and dynamically changing distribution of overall sales data. Traditional sales forecasting methods often struggle to adequately depict the changing patterns under different business scenarios when faced with complex sales data. This leads to insufficient stability and adaptability of forecast results in complex business environments, making it difficult to meet the more complex forecasting needs of actual business operations. Summary of the Invention
[0003] This application provides a training method, apparatus, electronic device, and storage medium for a sales forecasting model, which can effectively improve the robustness and scenario adaptability of the sales forecasting model, achieving more refined and accurate product sales forecasting. The above technical solution is as follows: Firstly, this application provides a method for training a sales forecasting model, including: Obtain a preset initial model and a preset time window set of sales data samples. The sales data sample set includes multiple sales data samples corresponding to multiple target products. Determine the sample feature information corresponding to each sales data sample. The sample feature information includes the sales volume statistical characteristics and sales time characteristics corresponding to the sales data sample. Based on the sample feature information of each sales data sample, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules, wherein the multiple preset expert modules are scenario-based expert sub-models in the preset initial model; The sales prediction model is obtained by training a preset initial model based on multiple sales data samples and target mapping relationships.
[0004] In one possible implementation, the target mapping relationship includes the correspondence between each sales data sample and multiple matching expert modules; Based on the sample feature information of each sales data sample, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules, including: Obtain scene feature information for each preset expert module; Based on the sample feature information of each sales data sample and the scene feature information of each preset expert module, the similarity information between each sales data sample and each preset expert module is determined. Based on the similarity information, multiple matching expert modules are identified for each sales data sample. Based on each sales data sample and its corresponding multiple matching expert modules, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules.
[0005] In one possible implementation, a pre-set initial model is trained based on multiple sales data samples and target mapping relationships to obtain a sales prediction model, including: Based on the target mapping relationship, determine the target sales data sample subset corresponding to each preset expert module. The target sales data sample subset includes multiple target sales data samples that match the preset expert module. Each preset expert module processes the corresponding subset of target sales data samples to obtain the expert prediction results output by each preset expert module. The expert prediction results are used to characterize multiple predicted sales information corresponding to multiple target sales data samples. Based on the prediction results of multiple experts output by multiple preset expert modules, a target loss function is constructed. The sales prediction model is obtained by training a preset initial model based on the objective loss function.
[0006] In one possible implementation, the target sales data sample subset corresponding to each preset expert module is determined based on the target mapping relationship, including: For each preset expert module, obtain multiple matching sales data samples corresponding to the preset expert module according to the target mapping relationship; Based on the expert processing rules corresponding to the preset expert module, multiple matching sales data samples are filtered to obtain multiple target sales data samples that match the preset expert module, and a subset of target sales data samples corresponding to the preset expert module is generated.
[0007] In one possible implementation, the target loss function includes a first loss function and / or a second loss function; The first loss function is used to characterize the intra-scenario prediction error between the expert prediction results output by each preset expert module for the corresponding target sales data sample subset and the actual sales label; the second loss function is used to characterize the inter-scenario consistency deviation between multiple expert prediction results output by multiple matching expert modules for the same target sales data sample.
[0008] In one possible implementation, when the target loss function includes a first loss function and a second loss function, the target loss function is constructed based on multiple expert prediction results output by multiple preset expert modules, including: Obtain multiple real sales tags corresponding to multiple target sales data samples; The first loss function is constructed based on multiple expert prediction results output by multiple preset expert modules and multiple real sales labels; A second loss function is constructed based on multiple expert prediction results output by multiple preset expert modules; The first loss function and the second loss function are weighted and combined to generate the target loss function.
[0009] In one possible implementation, a first loss function is constructed based on multiple expert predictions output by multiple preset expert modules and multiple real sales labels, including: For each target sales data sample, multiple matching predicted sales information corresponding to the target sales data sample is obtained from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample. Multiple matching and predicted sales information are fused to generate target prediction results corresponding to target sales data samples, thereby obtaining multiple target prediction results corresponding to multiple target sales data samples; The first loss function is constructed based on multiple real sales labels and multiple target prediction results.
[0010] In one possible implementation, multiple matching predicted sales information are fused to generate a target prediction result corresponding to the target sales data sample, including: Determine the expert weights of multiple matching expert modules corresponding to the target sales data sample; For each target sales data sample, multiple matching predicted sales information are weighted and fused based on multiple expert weights to obtain the target prediction result corresponding to the target sales data sample.
[0011] In one possible implementation, a second loss function is constructed based on multiple expert predictions output by multiple preset expert modules, including: For each target sales data sample, multiple matching predicted sales information corresponding to the target sales data sample is obtained from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample. Based on multiple matching predicted sales information, inter-scenario prediction deviation information is generated. The inter-scenario prediction deviation information is used to indicate the degree of difference between any two matching predicted sales information for the same target sales data sample. A second loss function is constructed based on the prediction bias information between scenarios.
[0012] In one possible implementation, each sales data sample represents the sales record information of the corresponding target product within each preset sales period. The sales record information includes: the sales date of the target product and the actual sales volume of the target product within the preset sales period. The preset sales period is a daily sales period, and the preset time window contains multiple preset sales periods. Determine the sample characteristic information corresponding to each sales data sample, including: Based on each sales data sample, construct the sales statistics features of the target product within a preset time window. The sales statistics features include at least one of the following: sales summation feature, sales extreme value feature, sales trend feature, and sales dispersion feature. Based on the sales date of the target product represented by each sales data sample, a sales time feature corresponding to each sales data sample is constructed. The sales time feature includes at least one of the following: basic time feature, periodic time feature, and special time feature. The sales statistics features and sales time features are optimized to obtain the sample feature information corresponding to each sales data sample.
[0013] In one possible implementation, the preset initial model also includes a gated network module and a fusion network module; The sales prediction model is obtained by training a pre-set initial model based on the objective loss function, including: Based on the objective loss function, the parameters of multiple preset expert modules, gated network modules and fusion network modules in the preset initial model are iteratively optimized until the trained preset initial model meets the preset convergence condition. Then, the trained preset initial model is determined as the sales prediction model.
[0014] In one possible implementation, multiple preset expert modules include a global expert module, which is used to process the sales data samples based on the sample feature information corresponding to each sales data sample in order to generate basic expert prediction results that do not depend on a specific scenario. The multiple preset expert modules also include at least one of the following: a promotional event scenario expert module, a new product scenario expert module, and an established product scenario expert module.
[0015] Secondly, this application provides a method for predicting commodity sales, including: Obtain historical sales data of the product to be predicted within a preset time window. The historical sales data includes sales record information of the product to be predicted in each preset sales period. Input historical sales data into the sales forecasting model to obtain the sales forecasting results of the product to be predicted. The aforementioned sales forecasting model is a sales forecasting model generated by a method provided by the first aspect or any possible implementation thereof.
[0016] Thirdly, this application provides a method for training a sales forecasting model, including: The first acquisition module is used to acquire a preset initial model and a sales data sample set within a preset time window. The sales data sample set includes multiple sales data samples corresponding to multiple target products. The first determining module is used to determine the sample feature information corresponding to each sales data sample. The sample feature information includes the sales volume statistical features and sales time features corresponding to the sales data samples. The construction module is used to build a target mapping relationship between multiple sales data samples and multiple preset expert modules based on the sample feature information of each sales data sample. The multiple preset expert modules are scenario-based expert sub-models in the preset initial model. The training module is used to train a preset initial model based on multiple sales data samples and target mapping relationships to obtain a sales prediction model.
[0017] Fourthly, this application provides a product sales forecasting device, comprising: The second acquisition module is used to acquire historical sales data of the product to be predicted within a preset time window. The historical sales data includes sales record information of the product to be predicted in each preset sales cycle. The second determining module is used to input historical sales data into the sales forecasting model to obtain the sales forecasting results of the product to be predicted. The aforementioned sales forecasting model is a sales forecasting model generated by a method provided by the first aspect or any possible implementation thereof.
[0018] Fifthly, this application provides an electronic device, including: a processor and a memory; wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method steps provided by the first aspect of this application or any possible implementation of the first aspect, or, when the processor executes the computer program, it implements the method steps provided by the second aspect of this application.
[0019] In a sixth aspect, this application provides a computer storage medium storing a plurality of instructions, which are adapted to be loaded by a processor and executed by the method steps provided by the first aspect of this application or any possible implementation thereof, or the instructions are adapted to be loaded by a processor and executed by the method steps provided by the second aspect of this application.
[0020] In a seventh aspect, this application provides a computer program product containing instructions that, when the computer program product is run on a computer or processor, causes the computer or processor to perform the method steps provided in the first aspect of this application or any possible implementation thereof, or causes the computer or processor to perform the method steps provided in the second aspect of this application.
[0021] This application obtains a pre-set initial model and a sales data sample set within a pre-set time window, where the sales data sample set includes multiple sales data samples corresponding to multiple target products. Then, it determines the sample feature information corresponding to each sales data sample, including sales volume statistics and sales time characteristics. Based on the sample feature information of each sales data sample, it constructs a target mapping relationship between multiple sales data samples and multiple pre-set expert modules, where the multiple pre-set expert modules are scenario-based expert sub-models in the pre-set initial model. The pre-set initial model is then trained according to the multiple sales data samples and the target mapping relationship to obtain a sales prediction model. Therefore, by matching historical sales data samples from multiple products and multiple periods with scenario-based expert sub-models, and by training the pre-set initial model specifically based on the mapping relationship, each pre-set expert module can fully learn specific sales patterns within its corresponding business scenario, significantly enhancing the model's fitting ability in different scenarios. Meanwhile, by constructing a target mapping relationship between sales data samples and preset expert modules, this application can achieve scenario-based splitting of input samples. This allows the originally mixed cross-scenario sales data to be accurately divided into more suitable expert modules for learning, thereby effectively reducing the interference caused by differences in sales patterns across different scenarios and improving the stability and convergence speed of the model training process. The final sales prediction model can not only accurately capture the sales trends and cyclical patterns of different products in different time periods, but also maintain high prediction accuracy and generalization ability when facing complex business conditions in different scenarios. Therefore, the embodiments of this application can effectively improve the robustness and scenario adaptability of the sales prediction model, achieving more refined and accurate product sales prediction. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 An exemplary system architecture diagram of a training method for a sales forecasting model provided in this application; Figure 2 A flowchart illustrating a training method for a sales forecasting model provided in this application; Figure 3 A flowchart illustrating a method for determining a target mapping relationship provided in this application; Figure 4 A flowchart illustrating a method for determining a sales forecasting model provided in this application; Figure 5A flowchart illustrating a method for constructing a target loss function provided in this application; Figure 6 A flowchart illustrating a method for constructing a first loss function provided in this application; Figure 7 A flowchart illustrating a method for constructing a second loss function provided in this application; Figure 8 A schematic diagram of the structure of a training device for a sales forecasting model provided in this application; Figure 9 A schematic diagram of a commodity sales forecasting device provided in this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0024] To make the features and advantages of this application more apparent and understandable, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] In the following description, when referring to the accompanying drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. Furthermore, in the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, in the description of this application, "multiple" refers to two or more.
[0026] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0028] In product sales scenarios, sales volume is influenced by various factors, resulting in a complex, diverse, and dynamically changing distribution of overall sales data. Traditional sales forecasting methods often struggle to adequately depict the changing patterns under different business scenarios when faced with complex sales data. This leads to insufficient stability and adaptability of forecast results in complex business environments, making it difficult to meet the more complex forecasting needs in actual business operations.
[0029] Therefore, this application provides a training method for a sales forecasting model to solve at least one of the above-mentioned technical problems in the prior art.
[0030] Please see Figure 1 , Figure 1 An exemplary system architecture diagram for a training method of a sales forecasting model provided in this application.
[0031] like Figure 1 As shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 serves as the medium for providing a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired or wireless communication links, such as wired communication links including fiber optic cables, twisted-pair cables, or coaxial cables, and wireless communication links including Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.
[0032] Terminal 101 can interact with server 103 via network 102 to receive messages from or send messages to server 103. Alternatively, terminal 101 can interact with server 103 via network 102 to receive messages or data sent to server 103 by other users. Terminal 101 can be hardware or software. When terminal 101 is hardware, it can be various electronic devices, including but not limited to tablet computers, laptops, and desktop computers. When terminal 101 is software, it can be installed in the aforementioned electronic devices and can be implemented as multiple software programs or software modules (e.g., to provide distributed services) or as a single software program or software module; no specific limitation is made here.
[0033] In this application, during the training preparation phase, terminal 101 first obtains a preset initial model and a sales data sample set within a preset time window from server 103 via network 102. The sales data sample set includes multiple sales data samples corresponding to multiple target products. After the samples and model architecture are prepared, terminal 101 determines the sample feature information corresponding to each sales data sample. The sample feature information includes the statistical features and time features corresponding to the sales data samples. Then, based on the sample feature information of each sales data sample, terminal 101 constructs a target mapping relationship between multiple sales data samples and multiple preset expert modules. The multiple preset expert modules are scenario-based expert sub-models in the preset initial model. Finally, the preset initial model is trained according to the multiple sales data samples and the target mapping relationship to obtain a sales prediction model.
[0034] Server 103 can be a server that provides various services. It should be noted that server 103 can be hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 103 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module; no specific limitations are made here.
[0035] Alternatively, the system architecture may not include server 103. In other words, server 103 may be an optional device in this application. That is, the method provided in this application can be applied to a system structure that only includes terminal 101. This application does not limit this.
[0036] It should be understood that Figure 1 The number of terminals, networks, and servers shown is only illustrative; the number can be any number of terminals, networks, and servers depending on the implementation requirements.
[0037] Please see Figure 2 , Figure 2 This is a flowchart illustrating a training method for a sales forecasting model provided in this application. The executing entity of this application can be a terminal executing the sales forecasting model training method, a processor within the terminal executing the sales forecasting model training method, or a sales forecasting model training service within the terminal executing the sales forecasting model training method. For ease of description, the following uses the processor within the terminal as an example to describe the specific execution process of the sales forecasting model training method.
[0038] Please see Figure 2 , Figure 2 This is a flowchart illustrating the training method for a sales forecasting model provided in this application. Figure 2 As shown, the training methods for sales forecasting models can include at least: S210, Obtain the preset initial model and the sales data sample set within the preset time window. The sales data sample set includes multiple sales data samples corresponding to multiple target products.
[0039] The preset initial model refers to the basic prediction model built before the training process begins. Specifically, it may include multiple preset expert modules (i.e., scenario-based expert sub-models) that can perform predictions for different business scenarios, different product types, or different sales regions. The preset initial model has not yet learned the target data distribution before training, and the model parameters are initially set for subsequent training and optimization based on sales data samples.
[0040] In one embodiment, the preset time window can be the time range used to construct the sales data sample set, specifically composed of multiple continuous or non-continuous preset sales cycles. This preset time window can be determined according to different business characteristics or forecasting needs. Specifically, the preset time window can include three categories: short-term time windows, medium-term time windows, and long-term time windows. The short-term time window can be the most recent 7 or 14 days, used to reflect the immediate sales changes of the target product within a short period; the medium-term time window can be set to the most recent 30 or 60 days, used to characterize the phased trends of the target product over a longer time span; and the long-term time window can be set to the most recent 90 or 180 days, used to depict the overall sales patterns of the target product over a longer period.
[0041] In one embodiment, the sales data sample set can be a collection of historical sales data collected within a preset time window, covering one or more sales regions and multiple target products. Specifically, the sales data sample set may include a sequence of sales data samples for each target product. This sequence includes sales record information for the target product within each preset sales period. The sales record information may include: the sales date of the target product and the actual sales volume of the target product within the preset sales period. The preset sales period is a daily sales period, such as a calendar day or a week.
[0042] Specifically, each sales data sample can contain daily sales records for the target product within a specific sales region. Each record can include the following fields: product identifier, sales date, actual sales volume, and sales channel. The product identifier can be a Stock Keeping Unit (SKU) code, representing the smallest available unit for inventory control. For example, assuming a preset time window of one month and a preset sales cycle of one calendar day, and the target product is product A (product identifier SKU12345), then for product A, a sales data sample corresponding to product A will be generated for each calendar day within one month. For instance, on November 1, 2025, the sales data sample generated for product A could be represented as: "SKU12345, 2025". 11 01, 28, directly operated store.
[0043] It should be noted that the aforementioned sales data sample set can be extracted from relevant business management tools or platforms. Taking a preset sales period of one natural day as an example, within each natural day, a sales data sample can be generated for each product in the business management tool or platform. These sales data samples generated on all natural days within the preset time window period then form the sales data sample sequence corresponding to each product. Furthermore, by aggregating the sales data sample sequences of multiple products within the preset time window, a sales data sample set covering multiple products and multiple sales regions can be obtained. This sales data sample set can comprehensively reflect the actual sales situation of each product in one or more sales regions under different channels and different time periods, providing a structured data foundation for subsequent feature construction, pattern mapping, and model training.
[0044] In one embodiment, the sales data sample set can also undergo data preprocessing. Specifically, firstly, for missing values in the sales data sample, the average sales volume of the next 30 days can be used to fill in the gaps, ensuring the continuity of the sales sequence of the target product within a preset time window and avoiding feature calculation bias caused by missing data. Secondly, abnormal sales data can be filtered out, for example, identifying abnormal zero sales records caused by system failures or abnormal order records exceeding 10 times the daily sales volume, and removing them from the sample to improve data reliability. Additionally, the sales data can be standardized, for example, using the Z-score standardization method to divide the original sales volume into (actual sales volume) and (actual sales volume) values. The transformation is performed using "monthly mean / monthly standard deviation" to map the sales data of target products for different SKUs to a unified numerical space, thereby providing a consistent data foundation for subsequent feature calculations and expert module matching.
[0045] S220, determine the sample feature information corresponding to each sales data sample. The sample feature information includes the sales volume statistics feature and sales time feature corresponding to the sales data sample.
[0046] Among them, sample feature information can be used to characterize the sales performance, fluctuation pattern, and temporal context semantics of the corresponding sales data sample within its preset time window.
[0047] Optionally, sales statistics features can be used to describe the sales volume, trend, and fluctuations of a target product within a preset time window, reflecting the sales performance of the target product over historical periods. Specifically, sales statistics features include at least one of the following feature fields: sales volume summation feature, sales volume extreme value feature, sales volume trend feature, and sales volume dispersion feature. The sales volume summation feature characterizes the overall sales volume of the target product within the preset time window, for example, it is calculated as the cumulative value of the sales volume of the product across all preset sales cycles within the preset time window; the sales volume extreme value feature reflects the sales volume of the target product in the maximum sales cycle (e.g., maximum single-day sales) and the minimum sales volume in the minimum sales cycle (e.g., minimum single-day sales) within the preset time window; the sales volume trend feature describes the changing trend of the target product's sales volume within the preset time window, and can be calculated based on linear regression slope, daily average sales growth rate, or segmented mean change rate, used to characterize the upward or downward trend of sales volume; the sales volume dispersion feature measures the degree of sales volume fluctuation, and can include indicators such as sales volume standard deviation, sales volume variance, and sales volume coefficient of variation, where the coefficient of variation is used to normalize the fluctuation differences between different sales volume levels.
[0048] Optionally, the sales time feature is used to characterize the semantic attributes of the date to which the sales data sample belongs on the time axis. Specifically, the sales time feature includes fields of at least one of the following features: basic time feature, periodic time feature, and special time feature. Among them, the basic time feature is used to characterize the basic attributes of the natural date to which the sales data sample belongs, including the product's listing time, the year, month, day, and day of the week corresponding to the sales date, etc.; the periodic time feature can be used to reflect the positional relationship of the sales date in the periodic pattern, including whether the sales date is a weekend, and whether it is in a periodic node such as a quarterly change month, which can help identify the periodic impact of weekly or quarterly cycles on sales; the special time feature can be used to characterize whether the sales date is at a special time node, including whether the sales date is a statutory holiday, 3 days before a holiday, 2 days after a holiday, whether it is a platform promotion date, a 30-day pre-promotion accumulation period, or a 5-day post-promotion return period, and may also include the promotion intensity, thereby describing the pulse-like sales changes caused by promotional activities or holidays.
[0049] S230: Based on the sample feature information of each sales data sample, construct the target mapping relationship between multiple sales data samples and multiple preset expert modules, wherein the multiple preset expert modules are scenario-based expert sub-models in the preset initial model.
[0050] In one embodiment, different preset expert modules among the multiple preset expert modules can be applied to different sales scenarios. For example, the multiple preset expert modules include a global expert module, which is used to process sales data samples based on the sample feature information corresponding to each sales data sample to generate basic expert prediction results that do not depend on a specific scenario. The multiple preset expert modules also include at least one of the following: a promotional event scenario expert module, a new product scenario expert module, and an established product scenario expert module. Among them, the global expert module can be used to process general target products that do not have significant scenario attributes, or to provide overall prediction support for complex products across scenarios. This expert module can learn the common distribution patterns of sales based on the full SKU data, serving as a supplement to other expert modules and improving the model's generalization ability in atypical scenarios. The Major Sales Promotion Scenario Expert Module is used to process target products during major sales promotion cycles across various sales channels. It models sales data samples exhibiting explosive growth and significant fluctuations within a short period. This module focuses on learning characteristics unique to major sales promotion scenarios, such as pulse-like sales patterns, add-to-cart behavior during the pre-sale period, operational strategy interventions, and delayed consumption during the post-sale period. This avoids interference from daily sales patterns in major sales promotion predictions, improving prediction accuracy during these scenarios. The New Product Scenario Expert Module is used to process target products in their early stages of launch and lacking long-term sales history. By analyzing the growth patterns of new products in the same category, market feedback data, and the ramp-up trends during the launch phase, it optimizes predictions in the early stages of a new product's lifecycle. This module effectively compensates for model bias caused by data sparsity in new products, improving prediction stability in cold-start scenarios. The Established Product Scenario Expert Module is suitable for target products with longer launch cycles and stable sales patterns. This module focuses on uncovering the cyclical fluctuations, seasonal patterns, and long-term average trends of established products over a longer time span, avoiding interference from complex scenario variables and improving prediction stability and robustness in established product scenarios. Optionally, multiple preset expert modules can also include special scenario expert modules. These modules are used to handle target products with special business characteristics, such as near-expiry clearance, customized products, and emergency products. These modules can combine the driving factors of special scenarios (such as price fluctuations, changes in promotional policies, and inventory pressure) with historical special case data to build models, thus avoiding the prediction failure of conventional models in abnormal scenarios.
[0051] Optionally, the target mapping relationship includes the correspondence between each sales data sample and multiple matching expert modules. Specifically, for each sales data sample, the corresponding multiple matching expert modules can be selected from multiple preset expert modules that have a high degree of matching with the current sales data sample in terms of scenario features. One sales data sample can correspond to multiple matching expert modules, thereby enabling the model to combine multiple scenario perspectives for prediction, improving the comprehensiveness and robustness of the prediction.
[0052] S240: Train the preset initial model based on multiple sales data samples and target mapping relationships to obtain a sales prediction model.
[0053] The sales forecasting model can be a deep forecasting model built upon multiple pre-defined expert modules. It can include multiple pre-defined expert modules for handling different sales scenarios, a gating network module for learning the relationship between samples and experts, and a fusion network module for outputting the final forecast result. This initial pre-defined model can learn the sales trend of the target product within a pre-defined sales period by jointly modeling the feature patterns, scenario attributes, and temporal regularities of sales data samples. After training, the resulting sales forecasting model can automatically select and activate suitable expert modules based on the historical sales data of the product to be predicted, outputting forecast results that better match the characteristics of the business scenario.
[0054] In one embodiment, the preset initial model further includes a gating network module and a fusion network module. The preset initial model is trained based on the objective loss function to obtain a sales prediction model, including: iteratively optimizing the parameters of multiple preset expert modules, gating network modules and fusion network modules in the preset initial model based on the objective loss function until the trained preset initial model meets the preset convergence condition, and then determining the trained preset initial model as the sales prediction model.
[0055] Among them, the preset convergence condition can be used to determine whether the training process of the preset initial model has stabilized. Specifically, it can include at least one of the following: the decrease of the target loss function is less than the preset magnitude threshold, the parameter change of each network module is lower than the preset change threshold in several consecutive iterations, and the number of iterations reaches the maximum number of training rounds.
[0056] Optionally, based on the aforementioned preset convergence condition, during the training process, it can be determined whether the current iteration meets the preset convergence condition. If so, the preset initial model is considered to have converged, and the current preset initial model is determined as the sales prediction model. This avoids computational waste and overfitting risks caused by overtraining, allowing the model to terminate training promptly after parameters stabilize, thereby improving training efficiency and enhancing the stability and generalization ability of the final sales prediction model.
[0057] In this embodiment, a sales data sample set within a preset time window is obtained from a preset initial model. This sales data sample set includes multiple sales data samples corresponding to multiple target products. Then, sample feature information corresponding to each sales data sample is determined. This sample feature information includes sales volume statistics and sales time characteristics corresponding to the sales data sample. Based on the sample feature information of each sales data sample, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules. These preset expert modules are scenario-based expert sub-models within the preset initial model. The preset initial model is then trained according to the multiple sales data samples and the target mapping relationship to obtain a sales prediction model. Therefore, by matching historical sales data samples from multiple products and multiple periods with scenario-based expert sub-models and training the preset initial model accordingly, each preset expert module can fully learn specific sales patterns within its corresponding business scenario, significantly enhancing the model's fitting ability in different scenarios. Meanwhile, by constructing a target mapping relationship between sales data samples and preset expert modules, this application can achieve scenario-based splitting of input samples. This allows the originally mixed cross-scenario sales data to be accurately divided into more suitable expert modules for learning, thereby effectively reducing the interference caused by differences in sales patterns across different scenarios and improving the stability and convergence speed of the model training process. The final sales prediction model can not only accurately capture the sales trends and cyclical patterns of different products in different time periods, but also maintain high prediction accuracy and generalization ability when facing complex business conditions in different scenarios. Therefore, the embodiments of this application can effectively improve the robustness and scenario adaptability of the sales prediction model, achieving more refined and accurate product sales prediction.
[0058] In one embodiment, in the above S220, determining the sample feature information corresponding to each sales data sample includes: constructing the sales volume statistical features of the target product within a preset time window based on each sales data sample; constructing the sales time features corresponding to each sales data sample based on the sales date of the target product represented by each sales data sample; and optimizing the sales volume statistical features and sales time features to obtain the sample feature information corresponding to each sales data sample.
[0059] Specifically, for each target product's sales data sample sequence, corresponding sales statistics characteristics can be calculated, such as sales summation characteristics, sales extreme value characteristics, sales trend characteristics, and sales dispersion characteristics. For example, for a target product, if the sales data in the most recent 30 days is {20, 32, 28, ...}, then the 30-day cumulative sales, 30-day maximum sales, 30-day average sales and their fluctuations can be calculated as the corresponding sales statistics feature fields for that target product.
[0060] In addition, the sales date corresponding to each sales data sample can be obtained first, and the corresponding basic time field, including year, month, day, and weekday information, can be parsed based on the sales date. Then, it is determined whether the sales date belongs to a preset periodic time identifier, such as whether it is a weekend or a month of quarterly transition, and the judgment result is recorded as a periodic time feature. Further, based on the sales date, a preset set of special dates is matched, including a set of statutory holidays, a set of dates for large-scale promotional activities, and their corresponding influence window ranges, to determine whether the sales date falls into any of the above special date intervals, and the matching result is recorded as a special time feature. Finally, the basic time feature, periodic time feature, and special time feature are combined to generate the sales time feature corresponding to the sales data sample.
[0061] Furthermore, in one embodiment, the feature fields in the sales statistics feature and the sales time feature can be filtered to optimize the sales statistics feature and the sales time feature.
[0062] Specifically, variance filtering can be performed first on sales statistics and sales time features. Variance filtering can be performed sequentially on each individual feature field within the sales statistics and sales time features. For example, based on all sales data samples within a preset time window, the variance of the value distribution of each feature field across all sales data samples can be calculated, and features with variances lower than a preset variance threshold can be filtered out. More specifically, for any feature field, the variance can be calculated based on the field's values across all sales data samples. If the feature field varies very little across different samples (variance lower than the preset variance threshold), it indicates that the feature field lacks discriminative power and can be directly removed.
[0063] Subsequently, correlation analysis can be performed on the remaining feature fields. Specifically, the Pearson correlation coefficient between any two feature fields can be calculated. When the absolute value of the correlation coefficient between two feature fields exceeds a preset correlation threshold (e.g., 0.8), it indicates a strong linear correlation between them, and the feature field can be considered a redundant feature pair. In this case, the correlation between each feature field and actual sales can be calculated, and features with a higher correlation to actual sales can be retained, while feature fields with lower correlation can be removed from the feature field set. For example, for the feature fields "30-day total sales" and "60-day total sales," if the absolute value of the correlation coefficient is greater than 0.8, then "30-day total sales," which has a higher correlation to actual sales, can be retained.
[0064] In one embodiment, the retained features can be further simplified based on feature importance ranking. Specifically, importance scores for each feature can be generated based on a tree model or a linear model, all features can be ranked by importance, and features with importance scores below a preset importance threshold can be removed, thereby obtaining the final feature set, which is the sample feature information corresponding to the sales data sample.
[0065] In this embodiment, the three-level screening mechanism of variance filtering, correlation analysis and importance ranking can effectively remove invalid features and highly redundant features, and retain the features that can most effectively reflect the sales change pattern of the target product, thereby providing more reliable and high-quality feature input for subsequent sample matching, expert module selection and model training.
[0066] Please see Figure 3 In S230 above, based on the sample feature information of each sales data sample, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules, which may include at least: S310: Obtain scene feature information from each preset expert module.
[0067] The scenario feature information can be used to characterize the business scenario attributes that the preset expert module is adapted to. Specifically, the scenario feature information may include feature description information used to distinguish different sales models, life cycle stages, promotional activity cycles, etc., and used to characterize the data patterns that the preset expert module is good at processing.
[0068] For example, the scenario feature information of the major promotion scenario expert module may include at least one field representing at least one of the following characteristics: whether it is in a major promotion cycle, whether it is in a major promotion accumulation period, whether it is in a return period, and a promotion intensity threshold, etc. The scenario feature information of the new product scenario expert module may include fields representing that the product's listing duration is less than a preset duration (e.g., 180 days), etc. The scenario feature information of the old product scenario expert module may include fields representing that the product's listing duration is not less than a preset duration (e.g., 180 days), etc. The scenario feature information of the special scenario expert module may include at least one field representing at least one of the following characteristics: near-expiry clearance, customized orders, and abnormal demand cycles, etc.
[0069] S320, based on the sample feature information of each sales data sample and the scene feature information of each preset expert module, determines the similarity information between each sales data sample and each preset expert module.
[0070] The similarity information is used to measure the degree of matching between the sample feature information of the sales data sample and the scene feature information of the preset expert module, and is used to select the preset expert module that is more suitable for processing the sales data sample, i.e., the matching expert module.
[0071] Optionally, the similarity information can be any of the following: Euclidean distance or Manhattan distance based on numerical features; matching scores based on categorical features (such as whether it is during a major promotion or a new product cycle); dynamic time-normalized distance based on sales trends; rule-based scoring; similarity prediction based on machine learning models, etc. The higher the similarity information, the more suitable the sales data sample is for processing by the corresponding preset expert module.
[0072] In one specific embodiment, for a sales data sample within a major promotional event period, the similarity score between the sales data sample and the major promotional expert module can be calculated based on the periodic time characteristics of the sample and the scene characteristic information of the major promotional expert module. At the same time, the similarity score between the sales data sample and the new product cold start expert module and the old product expert module can be calculated based on the launch time and other characteristics of the sales data sample, thereby obtaining multiple similarity information between the sales data sample and each preset expert module.
[0073] S330, based on the similarity information, determines multiple matching expert modules for each sales data sample.
[0074] Furthermore, based on the obtained similarity information, multiple preset expert modules with high similarity information can be selected as the matching expert modules corresponding to each sales data sample. For example, for each sales data sample, the preset expert modules whose similarity information to the corresponding sales data sample exceeds a preset similarity threshold can be determined as the corresponding matching expert modules; alternatively, the global expert module can be directly determined as the matching expert module corresponding to each sales data sample, thereby obtaining multiple matching expert modules corresponding to each sales data sample.
[0075] S340: Based on each sales data sample and its corresponding multiple matching expert modules, construct a target mapping relationship between multiple sales data samples and multiple preset expert modules.
[0076] Specifically, the above steps yield the correspondence between "sales data samples" and "multiple matching expert modules." Each sales data sample is then bound to its matched multiple preset expert modules (i.e., matching expert modules), generating a target mapping relationship. This mapping relationship can be used to segment training samples into scenarios during the training phase: each preset expert module only receives a subset of target sales data samples that match its scenario features, thereby achieving scenario-oriented training of the preset expert modules and avoiding noise interference from cross-scenario samples to the model. Ultimately, this target mapping relationship can serve as the input structure for the preset initial model training process, enabling the sales prediction model to achieve structured, scenario-specific joint learning during the training phase.
[0077] In this embodiment, by automatically dividing a large number of historical sales samples into different preset expert modules according to their own characteristics and scene characteristics, sample diversion is achieved, thereby enabling refined scene identification of samples. This avoids interference from erroneous scene samples with model learning and training, effectively improving the prediction robustness of difficult and special samples. Ultimately, it provides structured expert input for subsequent model training, enabling the obtained sales prediction model to have better prediction performance under cross-scene, multi-category, and multi-region conditions.
[0078] Please see Figure 4 In step S240 above, the sales prediction model is obtained by training a preset initial model based on multiple sales data samples and target mapping relationships, which may include at least: S410, determine the target sales data sample subset corresponding to each preset expert module according to the target mapping relationship. The target sales data sample subset includes multiple target sales data samples that match the preset expert module.
[0079] In one embodiment, the target sales data sample subset corresponding to each preset expert module is used to represent a set of multiple sales data samples that are determined to match the preset expert module in the target mapping relationship. For example, the target sales data sample subset corresponding to the promotional event expert module includes sales data samples during the promotional event, the target sales data sample subset corresponding to the new product scenario expert module includes sales data samples whose launch time has not reached the preset time, the target sales data sample subset corresponding to the old product scenario expert module includes sales data samples whose launch time has reached the preset time, and the target sales data sample subset corresponding to the global expert module includes sales data samples from all scenarios.
[0080] In another embodiment, the subset of target sales data samples corresponding to each preset expert module can also be used to represent a set of higher-quality training samples that satisfy both the similarity information in the target mapping relationship and the scene feature constraints in the expert processing rules.
[0081] Specifically, the sales data samples corresponding to the target mapping relationship for each preset expert module can be further filtered. Specifically, in S410, determining the subset of target sales data samples corresponding to each preset expert module based on the target mapping relationship can include: for each preset expert module, obtaining multiple matching sales data samples corresponding to the preset expert module according to the target mapping relationship; filtering the multiple matching sales data samples based on the expert processing rules corresponding to the preset expert module to obtain multiple target sales data samples that match the preset expert module, thus generating the subset of target sales data samples corresponding to the preset expert module.
[0082] The "matched sales data sample" refers to the set of sales data samples initially selected for a pre-defined expert module based on the target mapping relationship. Sales data samples in this set are considered suitable candidate samples for processing by the expert module because they have a high degree of similarity to the scenario feature information of the pre-defined expert module.
[0083] Optionally, expert processing rules represent the sample validity constraints pre-set by each preset expert module for its own business scenario, used to further filter the matched sales data samples. Expert processing rules can be defined by the core features of the scenarios in which the experts are proficient, used to remove noise and enhance the scenario of the initially matched samples.
[0084] Specifically, the expert processing rules for the "Major Promotion Scenario Expert Module" can require that the matched sales data samples meet the following conditions: the sales date of the target product included in the matched sales data sample falls within the major promotion activity period, and the promotional intensity corresponding to the matched sales data sample is not lower than a preset promotional intensity threshold. This rule can filter out target sales data samples that reflect the short-term surge and drastic fluctuations in sales during the major promotion period. The expert processing rules for the "New Product Scenario Expert Module" can require that the target product corresponding to the matched sales data sample has not reached a preset launch duration threshold (e.g., launch duration is less than the preset duration), and that the data records of the matched sales data sample remain continuous within the preset time window, without any interruption in sales data for consecutive preset sales cycles. These rules can eliminate matched sales data samples that no longer reflect the characteristics of the new product's growth stage due to missing data or prolonged launch, thus retaining high-quality target sales data samples that reflect the early growth trend of the new product. The expert processing rules for the "Old Product Stability Expert Module" can be used to filter out target sales data samples that reflect the sales characteristics of products in a mature and stable stage. This processing rule can require that the product corresponding to the matching sales data sample has been on the market for a preset duration, and that the fluctuation of the target product's sales volume represented by the matching sales data sample within a preset time window does not exceed a preset fluctuation threshold. This allows the selected target sales data sample to reflect the relatively stable and predictable sales characteristics of established products during their stable maturity period. Furthermore, for the global expert module, the corresponding expert processing rule can only require that the data integrity of the sales data sample within a preset time window meets a preset standard, such as the percentage of missing values not exceeding a preset proportion.
[0085] In one embodiment, multiple matching sales data samples can be input into the expert processing rules of the corresponding preset expert module for filtering, thereby obtaining multiple target sales data samples that are more suitable for the preset expert module, thus generating a subset of target sales data samples corresponding to the preset expert module.
[0086] In this embodiment, by introducing scenario-feature-based expert processing rules for each preset expert module, the initially matched sales data samples are further filtered. This ensures that each preset expert module only receives a subset of target sales data samples highly relevant to its business scenario, effectively eliminating cross-scenario samples, abnormal samples, and noisy samples. Consequently, each expert module can focus on its preferred data patterns during the training phase, improving the learning accuracy and stability of the scenario model. Simultaneously, it reduces the interference of data noise on model parameters, thereby improving the overall sales prediction model's scenario adaptability, prediction accuracy, and generalization ability.
[0087] S420 processes the corresponding subset of target sales data samples through each preset expert module to obtain the expert prediction results output by each preset expert module. The expert prediction results are used to characterize multiple predicted sales information corresponding to multiple target sales data samples.
[0088] In one embodiment, each preset expert module in the preset initial model processes its corresponding subset of target sales data samples to output an expert prediction result representing sales forecast. The predicted sales information represented by the expert prediction results output by each preset expert module corresponds one-to-one with the target sales data samples in the subset of target sales data samples.
[0089] For example, the expert module for major sales events can receive 200 target sales data samples, and output a predicted sales volume for each target sales data sample. That is, the expert prediction results output by the expert module for major sales events contain 200 predicted sales volume information.
[0090] S430 constructs the target loss function based on multiple expert prediction results output by multiple preset expert modules.
[0091] In one embodiment, the target loss function required for model training can be constructed based on all expert prediction results output by all preset expert modules.
[0092] Optionally, the target loss function includes a first loss function and / or a second loss function; wherein, the first loss function is used to characterize the intra-scenario prediction error between the expert prediction results output by each preset expert module for the corresponding target sales data sample subset and the actual sales label; the second loss function is used to characterize the inter-scenario consistency deviation between multiple expert prediction results output by multiple matching expert modules for the same target sales data sample.
[0093] Specifically, when the target loss function includes a first loss function and a second loss function, the target loss function can be a weighted sum of the first loss function and the second loss function.
[0094] S440 trains a pre-set initial model based on the objective loss function to obtain a sales prediction model.
[0095] Specifically, gradient descent or other optimization algorithms can be used to continuously update the parameters of the preset initial model based on the target loss function, ultimately obtaining a sales prediction model after training convergence.
[0096] In this embodiment, firstly, a subset of target sales data samples corresponding to each preset expert module is determined based on scene features and sample features. This ensures that each expert module only receives high-quality training samples that it is best at processing, avoiding interference from cross-scene data noise on the model. Subsequently, each expert module independently models its target sales data sample subset, outputting scenario-based expert prediction results. Then, based on all expert prediction results, a target loss function with fusion consistency and scenario-specific accuracy constraints is constructed, achieving cross-scene joint optimization. Finally, a preset initial model is trained using this target loss function, enabling the model to possess higher prediction accuracy, stability, and generalization ability in different business scenarios such as promotional events, new product launches, and existing product launches. Thus, this application effectively improves the overall performance of the sales prediction model through scenario-specific sample selection, expert-oriented modeling, and multi-loss collaborative constraints.
[0097] Please see Figure 5 In one embodiment, when the target loss function includes a first loss function and a second loss function, in the above-described S430, the target loss function is constructed based on multiple expert prediction results output by multiple preset expert modules, including: S510 retrieves multiple real sales labels corresponding to multiple target sales data samples.
[0098] Each target sales data sample corresponds one-to-one with a real sales label, and the real sales label corresponding to each target sales data sample can be the actual sales volume of the target product contained in that target sales data sample.
[0099] S520 constructs a first loss function based on multiple expert prediction results output by multiple preset expert modules and multiple real sales labels.
[0100] Specifically, the error between the predicted results of each target sales data sample and the actual sales label can be measured, and the errors of all samples can be summarized and calculated to generate a first loss function to characterize the prediction bias within the scenario.
[0101] Optionally, any one of the loss forms, such as mean squared error, root mean square error, or mean absolute error, can be used to quantify the in-scenario prediction error between the expert prediction results output by each preset expert module for the corresponding target sales data sample subset and the actual sales label.
[0102] S530 constructs a second loss function based on multiple expert prediction results output by multiple preset expert modules.
[0103] Specifically, for the same target sales data sample, the difference measure between any two prediction results can be calculated from the expert prediction results output by multiple matching expert modules, and the difference results of all samples can be summarized and calculated to generate a second loss function to characterize the consistency deviation between scenarios.
[0104] Optionally, the absolute value of the difference between expert prediction results, the squared difference, or the difference measure based on the proportional deviation can be used to quantify the inter-scenario consistency error between multiple expert prediction results output by multiple matching expert modules for the same target sales data sample, thereby constraining the prediction differences of different scenario expert modules on the same sample to remain within the range allowed by business logic.
[0105] S540, the first loss function and the second loss function are weighted and combined to generate the target loss function.
[0106] In one embodiment, the target loss function can be a linear weighted sum of a first loss function and a second loss function. For example, the first and second loss functions can be weighted and combined using preset weight parameters α and β (the sum of α and β is 1), so that the target loss function can improve the prediction accuracy within the scene while taking into account the prediction coordination across expert modules. Specifically, the target loss function L can be expressed as: L = α × L1 + β × L2, where L1 is the first loss function and L2 is the second loss function.
[0107] It should be noted that the specific values of α and β can be flexibly set according to the focus of the model training task, so that the first loss function and the second loss function can work together in the overall optimization process. For example, the weight of the first loss function can be set to a higher value, such as 0.7, and the weight of the second loss function can be set to a lower value, such as 0.3. This allows the importance of prediction accuracy within a scene to be given a higher weight, while the consistency between scenes is used as an auxiliary optimization objective, thus balancing the model's specialized learning with the need for global consistency.
[0108] In this embodiment, by simultaneously introducing a first loss function and a second loss function into the target loss function and using a weighted combination for unified optimization, the model training phase can balance the fine-fit capability within a given scenario with consistency constraints across different expert scenarios. Through this loss combination mechanism, the model can achieve more stable, reasonable, and business logic-consistent sales prediction results in different scenarios, significantly improving the accuracy and robustness of the final sales prediction model.
[0109] Please see Figure 6In S520, a first loss function is constructed based on multiple expert prediction results output by multiple preset expert modules and multiple real sales labels, including: S610: For each target sales data sample, obtain multiple matching predicted sales information corresponding to the target sales data sample from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample.
[0110] Among them, multiple matching predicted sales information can be used to characterize multiple sets of predicted values of the same target sales data sample from the perspective of different scenario-based expert sub-models.
[0111] S620 integrates multiple matching and predicted sales information to generate target prediction results corresponding to target sales data samples, thereby obtaining multiple target prediction results corresponding to multiple target sales data samples.
[0112] Specifically, after acquiring multiple matching predicted sales information, these multiple matching predicted sales information can be merged to generate a unique prediction result corresponding to the target sales data sample, i.e., the target prediction result. The target prediction result can represent the final predicted value given by the model for the target sales data sample after integrating the prediction capabilities of multiple matching expert modules.
[0113] For example, fusing multiple matched predicted sales information to generate a target prediction result corresponding to a target sales data sample may include: determining multiple expert weights of multiple matching expert modules corresponding to the target sales data sample; and for each target sales data sample, performing weighted fusion of multiple matched predicted sales information based on multiple expert weights to obtain the target prediction result corresponding to the target sales data sample.
[0114] Among them, the matching expert module and the expert weight are in one-to-one correspondence. The expert weight can be used to characterize the relative contribution of the corresponding matching expert module in generating the target prediction result corresponding to the target sales data sample.
[0115] In one embodiment, the process of weighted fusion of multiple matched predicted sales information based on multiple expert weights can be achieved through the collaborative operation of the gated network module and the fusion network module in the preset initial model.
[0116] Specifically, the gating network module generates dynamically adjusted expert weights for multiple matching expert modules based on the sample feature information of the target sales data sample. For example, the gating network module can employ a multilayer perceptron structure to encode the aforementioned high-dimensional sample feature information and output multiple expert weights through a normalized exponential function (Softmax), ensuring that all matching expert weights are non-negative and their sum is 1. In a specific embodiment, when the target sales data sample is within a major promotional activity period, the gating network module can capture the promotional scene features of the sample, thereby assigning higher weights to the promotional scene expert module, for example, increasing the weight to 0.4–0.6; when the target product corresponding to the target sales data sample has a short listing period and exhibits obvious cold-start characteristics, the gating network module can automatically increase the weight of the new product scene expert module, making the new product scene expert module dominant in the weighted fusion of this sample; for sales data samples lacking clear scene attributes, the gating network module can assign higher weights to the global expert module to improve the robustness of the prediction results.
[0117] Furthermore, after obtaining the expert weights corresponding to multiple matching expert modules, the fusion network module can perform weighted fusion of multiple matched predicted sales information based on the expert weights. Specifically, the fusion network module can weight the predicted sales information output by each matching expert module with the corresponding expert weight one by one, and summarize the weighted results to obtain the target prediction result corresponding to the target sales data sample.
[0118] In a further embodiment, the fusion network module can also adjust the weighting process based on the expert weights and prediction confidence information of each matching expert module, so that the matching expert module with higher prediction confidence has a greater influence on the final prediction value. The prediction confidence can be calculated based on factors such as the prediction accuracy of the expert module on historical samples and the scene adaptability of the samples. Specifically, the target prediction result can be obtained using the following formula: Y final =Σ(W i ×Y i ×S i ); Among them, Y final W represents the target prediction result. i Y represents the expert weight of the i-th matching expert module. i This indicates that the matching expert module outputs predicted sales information based on the target sales data sample, S i This represents the real-time prediction confidence of the matching expert module. Optionally, the expert confidence S... iThe prediction accuracy can be determined by combining the prediction accuracy of the matching expert module on historical samples with the fit between the current target sales data sample and the scenario characteristics. By introducing expert confidence, the prediction results can be made more reliant on the matching expert module, which has higher reliability in the current scenario, thereby further improving the stability and accuracy of the target prediction results.
[0119] In another embodiment, the fusion process may further include a prediction consistency verification step to enhance the global reasonableness of the final prediction result. Specifically, when the mean deviation of the predicted sales information from a certain matching expert module from the predicted sales information output by other matching expert modules exceeds a preset deviation threshold (e.g., 30%), a weight backtracking mechanism can be triggered. The gating network module readjusts the weight of the matching expert module, automatically reducing its influence on the final target prediction result and preventing distortion of the target prediction result due to abnormal predictions.
[0120] For example, suppose a target sales data sample corresponds to three matching expert modules, and these three modules output predicted sales information of 20, 24, and 30, respectively. The gating network module dynamically generates three expert weights of 0.5, 0.3, and 0.2 based on the sample features. The fusion network module can then weight the three predicted sales information according to these expert weights to obtain the target prediction result: 20 × 0.5 + 24 × 0.3 + 30 × 0.2 = 23.8. In an embodiment that further considers prediction confidence, the fusion network module can also perform a secondary adjustment on the weighted result based on the confidence level of each matching expert module, making the final prediction result more reliable.
[0121] By using the dynamic attention weight allocation mechanism and weighted fusion strategy described above, the model can automatically select the more suitable expert module in different business scenarios. This not only effectively improves the accuracy of sample-level predictions but also reduces the prediction bias caused by traditional fixed-weight methods, thereby improving the stability, rationality, and generalization ability of the final results.
[0122] S630 constructs a first loss function based on multiple real sales labels and multiple target prediction results.
[0123] Specifically, for each target sales data sample, the error between the target prediction result output by the fusion network module and the corresponding actual sales label can be measured to obtain the prediction deviation information for each sample. Subsequently, the prediction deviation information for all target sales data samples is aggregated and calculated to generate the first loss function. For example, the first loss function can be constructed using the root mean square error (RMSE) form, and its specific calculation method can be expressed as follows:
[0124] Where L1 represents the first loss function, Y pred Y represents the target prediction result corresponding to the target sales data sample. true This represents the actual sales volume label for the sample, and N represents the number of target sales data samples used in the calculation.
[0125] In this embodiment, a first loss function is introduced to measure the deviation between the target prediction result and the actual sales label. The prediction error of all samples is aggregated using methods such as root mean square error. This effectively enhances the model's fitting ability in various sub-scenarios, allowing each preset expert module to obtain higher-precision supervisory feedback on its preferred data distribution. Furthermore, this error measurement method significantly penalizes samples with large prediction deviations, thereby promoting automatic calibration of unstable prediction patterns during model training and improving the overall model's training convergence efficiency and prediction stability.
[0126] Please see Figure 7 In one embodiment, in S530, a second loss function is constructed based on multiple expert prediction results output by multiple preset expert modules, including: S710: For each target sales data sample, obtain multiple matching predicted sales information corresponding to the target sales data sample from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample.
[0127] Among them, multiple matching predicted sales information can be used to characterize multiple sets of predicted values of the same target sales data sample from the perspective of different scenario-based expert sub-models.
[0128] Specifically, for each target sales data sample, multiple matched predicted sales information corresponding to that target sales data sample can be extracted one by one from the prediction results of multiple matching expert modules. Each matching expert module outputs a set of predicted values based on its independent scenario-based modeling capabilities. Therefore, for the same target sales data sample, two or more sets of predicted sales information can be obtained to reflect the prediction opinions of different experts on that sample.
[0129] For example, for a target sales data sample within a major promotional period, the major promotional scenario expert module, the new product scenario expert module, and the global expert module may all generate prediction results, such as 120, 95, and 110, which constitute multiple matching predicted sales information corresponding to the sample.
[0130] S720 generates inter-scenario prediction deviation information based on multiple matched predicted sales information. The inter-scenario prediction deviation information is used to indicate the degree of difference between any two matched predicted sales information for the same target sales data sample.
[0131] The inter-scenario prediction deviation information may include a deviation penalty result obtained by comparing the relative deviation value between any two matched predicted sales data for each target sales data sample with a preset scenario logic threshold.
[0132] Specifically, it can predict sales information Y for any two matched samples of sales data targeting the same goal. i With Y j The set of deviation metrics is used to characterize whether the prediction differences between different matching expert modules on this sample exceed a preset scene logic threshold. This deviation metric can be expressed as:
[0133] Where T is the preset scene logic threshold.
[0134] S730 constructs a second loss function based on the prediction deviation information between scenes.
[0135] Optionally, after obtaining all the inter-scenario prediction deviation terms corresponding to each target sales data sample, the deviation terms of each sample can be summed or averaged, and the inter-scenario prediction deviation information of all target sales data samples can be summarized to construct a second loss function.
[0136] Specifically, the second loss function can be expressed as:
[0137] Where P represents the set of pairwise matching expert combinations consisting of multiple matching expert modules for the same target sales data sample; Y i and Y j These represent the predicted sales information output by the two matching expert modules mentioned above; T is a preset scenario logic threshold, used to indicate the upper limit of the prediction difference between different scenario expert modules within the allowable range of business logic. When the relative deviation ratio between any two predicted sales information does not exceed the threshold T, the corresponding item is set to 0; when the deviation exceeds the threshold T, a positive penalty value is generated, thereby constraining the prediction consistency of different scenario expert modules on the same target sales data sample and avoiding prediction conflicts between scenarios that violate business rules.
[0138] This application's embodiments can constrain the consistency of prediction results output by multiple matching expert modules for the same target sales data sample across different scenarios, effectively avoiding excessive dispersion or contradiction in predictions from experts in different scenarios on the same sample. By measuring the relative deviation of pairwise matched sales prediction information and generating a penalty term based on whether the deviation exceeds a preset scenario logic threshold, unreasonable differences in cross-scenario expert prediction results can be actively suppressed during training. This enhances the model's collaborative ability across different business scenarios, making the final sales prediction results more consistent with business logic, more stable, and more reliable.
[0139] This application also provides a method for predicting commodity sales, including: acquiring historical sales data of the commodity to be predicted within a preset time window, the historical sales data including sales record information of the commodity to be predicted in each preset sales cycle; inputting the historical sales data into a sales prediction model to obtain the commodity sales prediction result output by the sales prediction model for the commodity to be predicted; wherein, the sales prediction model is a sales prediction model generated by the training method of the sales prediction model.
[0140] In one embodiment, the product sales forecasting method can be applied to product sales forecasting scenarios on e-commerce platforms. Specifically, historical sales data of the product to be predicted within a preset time window can be obtained first. This historical sales data includes the actual sales volume and related sales records of the product within a preset sales cycle, such as daily or weekly. Subsequently, the historical sales data is input into the sales forecasting model obtained through the aforementioned training method. The model can then infer the future sales volume of the product to be predicted based on its internal scenario-based expert module, gating network module, and fusion network module, outputting the corresponding product sales forecast result. Thus, by applying the trained sales forecasting model to the actual forecasting process, corresponding sales forecast results can be quickly generated after inputting historical sales data from real business scenarios, resulting in simple deployment and high forecasting efficiency. Since the forecasting model has completed scenario-based expert learning, dynamic weight allocation, and multi-expert result fusion during the training phase, it can automatically adapt to different product sales scenarios during actual forecasting, generating more stable and accurate forecast results, thereby effectively supporting business needs such as inventory management, procurement decisions, and operational strategy formulation.
[0141] In one embodiment, the sales forecasting model can support outputting sales forecast results for multiple preset time windows for different business needs, including short-term forecast windows, medium-term forecast windows, and long-term forecast windows. Specifically, the sales forecasting model can not only output a single-point forecast value for a specific future sales cycle (e.g., the next day), but also flexibly generate sequential forecast results for multiple consecutive future sales cycles (e.g., the next 7 days, 14 days, or 30 days) in a single inference process.
[0142] It's important to note that during the model training phase, it's unnecessary to build separate models for different preset time windows. Specifically, multiple prediction targets (such as short-term prediction labels, medium-term prediction labels, and long-term prediction labels) can be introduced simultaneously during the same training process. This allows the model to learn sales variation patterns across different prediction spans within a unified parameter space. Consequently, after a single training iteration, the model can selectively output a single predicted value or a multi-day prediction sequence based on business needs. For example, when the business only requires short-term sales forecasts, the model can directly output predictions for the next 1 day or 7 days; when the business needs medium-term planning or inventory optimization, the same model can output daily prediction sequences for the next 30 days, achieving flexible adaptation of the prediction window.
[0143] The embodiments of this application not only reduce the computational resources and maintenance costs of training multiple models for different preset time windows, but also improve the uniformity and scalability of the prediction system, enabling the sales prediction model to flexibly adjust the output structure according to the real-time needs of the business side, thereby enhancing the applicability of the model in actual operation scenarios.
[0144] In one embodiment, the preset initial model may also include an output and verification module. The output and verification module can be used to perform multi-dimensional output and accuracy enhancement processing on the sales forecast results of the product after the sales inference of the product to be predicted is completed in the inference stage. It can also output auxiliary decision-making information at the same time, so that relevant business personnel can further evaluate the reliability and influencing factors of the forecast results based on the sales forecast results.
[0145] Optionally, the aforementioned auxiliary decision-making information may include at least one of the following: expert weight information of each preset expert module, pre-set reliability information, and key influencing factors. The expert weight information can be generated based on the output of the gating network module. The prediction confidence information can be generated based on the model's historical prediction error or the prediction stability of the expert module. Specifically, the system can call historical prediction error statistics, use the root mean square error (RMSE) or other error indicators to calculate the confidence interval of the prediction result, and output this confidence interval synchronously with the prediction result to characterize the confidence boundary of the prediction value. The key influencing factor information can represent the feature fields that have a significant impact on the final sales prediction result. Specifically, it can determine at least some of the features that have the greatest impact on the current sales prediction result through feature contribution calculation methods, and output these features and their contribution ratios to demonstrate the main influencing factors of the prediction result. The feature contribution calculation method may include: methods for interpreting machine learning model predictions (Shapley Additive exPlanations, SHAP) or random forest feature contribution analysis, etc.
[0146] For example, in the process of generating sales forecasts for a product to be predicted, the output and validation module can generate the following auxiliary decision-making information: the final predicted value for the product is 100 units; the expert weights of the expert modules for the promotional scenario, the old product scenario, and the global expert module generated by the gating network module for this sample are 0.45, 0.30, and 0.25, respectively; the prediction confidence interval is calculated to be [85, 115] based on the historical RMSE=7.5; and the main factors affecting the prediction result are identified through feature contribution analysis, including the feature fields "promotional intensity," "sales growth rate in the 14 days before the promotion," and "day of the week," with contribution percentages of 35%, 28%, and 12%, respectively. Based on the above information, relevant business personnel can clearly determine the main expert modules referenced in this prediction, the confidence range of the predicted value, and the core factors affecting the prediction.
[0147] Through the above implementation methods, the output and verification module can provide interpretability enhancement processing for sales forecast results, so that the forecast results not only include core forecast values, but also key auxiliary decision-making information, thereby improving the transparency and credibility of the forecast results and the usability of the model on the business side, avoiding the technical problem that traditional models are difficult to be effectively applied due to a lack of interpretability.
[0148] In one embodiment, the preset initial model may further include a verification and iteration module, which is used to perform rolling verification of the prediction accuracy of the model based on real business data after the sales prediction model inference is completed, and to trigger automatic iterative updates of the sales prediction model when the accuracy does not meet the preset accuracy conditions, thereby constructing a closed-loop optimization mechanism so that the sales prediction model can maintain a high prediction accuracy when the sales scenario of the target product changes dynamically.
[0149] Specifically, in one embodiment, the accuracy verification and iteration module can first evaluate the predictive performance of the sales forecasting model in real time using a rolling verification method. For example, after each forecasting cycle (e.g., 30 days), the actual sales volume for each day within that forecasting cycle can be compared with the corresponding sales forecast result output by the model. During the rolling verification process, the predictive performance of the model can be quantified based on preset accuracy evaluation metrics. Accuracy evaluation metrics may include Mean Absolute Error (MAE) and / or Mean Absolute Percent Error (MAPE).
[0150] Specifically, the mean absolute error (MAE) can be calculated using the following formula:
[0151] Specifically, the Mean Absolute Percentage Error (MAPE) can be calculated using the following formula:
[0152] in, Indicates the actual sales volume of the product. This represents the predicted sales volume of the product, where n represents the number of historical sales data points used in the inference.
[0153] Furthermore, a MAPE ≤ 25% can be considered as indicating high prediction accuracy in this scenario, while 25% < MAPE ≤ 35% can be considered as acceptable prediction accuracy. Additionally, when the MAPE exceeds 35%, the model's intelligent iteration mechanism can be automatically triggered to adaptively update key components, addressing the decline in prediction accuracy caused by changes in the target product scenario. Specifically, the sales prediction model can first re-optimize the feature selection rules, re-executing the process of selecting each feature field in sales statistics and sales time features to remove outdated feature fields (e.g., a target product has transitioned from a new product stage to an established product stage, rendering new product features inapplicable) and supplementing with new feature fields that reflect the current sales scenario. Furthermore, the expert weights of each preset expert module can be recalculated. Moreover, the multilayer perceptron parameters and attention mechanism weights in the gating network module can be updated to enable the gating network to more accurately identify key features in the current scenario, thereby improving the rationality of weight allocation in the weighted fusion process of the corresponding expert modules.
[0154] In this embodiment, the prediction accuracy verification and iteration module enables the model to dynamically and adaptively adjust throughout the entire lifecycle of the target product, including different sales scenarios such as new product stage, old product stage, changes in promotional cycle, switching of promotional strategies, and near-expiry clearance, thereby continuously maintaining the stability and accuracy of the prediction results and providing long-term and effective support for subsequent business decisions.
[0155] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0156] The inventive concept based on the training method of the above sales forecasting model, such as... Figure 8As shown, this application also provides a training apparatus 800 for a sales forecasting model to implement the training method of the sales forecasting model involved above. The training apparatus 800 for the sales forecasting model includes: The first acquisition module 810 is used to acquire a preset initial model and a sales data sample set within a preset time window. The sales data sample set includes multiple sales data samples corresponding to multiple target products. The first determining module 820 is used to determine the sample feature information corresponding to each sales data sample. The sample feature information includes the sales volume statistical features and sales time features corresponding to the sales data samples. Module 830 is used to construct a target mapping relationship between multiple sales data samples and multiple preset expert modules based on the sample feature information of each sales data sample. The multiple preset expert modules are scenario-based expert sub-models in the preset initial model. Training module 840 is used to train a preset initial model based on multiple sales data samples and target mapping relationships to obtain a sales prediction model.
[0157] In one possible implementation, the target mapping relationship includes the correspondence between each sales data sample and multiple matching expert modules; the construction module 830 is specifically used for: Obtain scene feature information for each preset expert module; Based on the sample feature information of each sales data sample and the scene feature information of each preset expert module, the similarity information between each sales data sample and each preset expert module is determined. Based on the similarity information, multiple matching expert modules are identified for each sales data sample. Based on each sales data sample and its corresponding multiple matching expert modules, a target mapping relationship is constructed between multiple sales data samples and multiple preset expert modules.
[0158] In one possible implementation, the training module 840 is specifically used for: Based on the target mapping relationship, determine the target sales data sample subset corresponding to each preset expert module. The target sales data sample subset includes multiple target sales data samples that match the preset expert module. Each preset expert module processes the corresponding subset of target sales data samples to obtain the expert prediction results output by each preset expert module. The expert prediction results are used to characterize multiple predicted sales information corresponding to multiple target sales data samples. Based on the prediction results of multiple experts output by multiple preset expert modules, a target loss function is constructed. The sales prediction model is obtained by training a preset initial model based on the objective loss function.
[0159] In one possible implementation, the training module 840 is specifically used for: For each preset expert module, obtain multiple matching sales data samples corresponding to the preset expert module according to the target mapping relationship; Based on the expert processing rules corresponding to the preset expert module, multiple matching sales data samples are filtered to obtain multiple target sales data samples that match the preset expert module, and a subset of target sales data samples corresponding to the preset expert module is generated.
[0160] In one possible implementation, the target loss function includes a first loss function and / or a second loss function; The first loss function is used to characterize the intra-scenario prediction error between the expert prediction results output by each preset expert module for the corresponding target sales data sample subset and the actual sales label; the second loss function is used to characterize the inter-scenario consistency deviation between multiple expert prediction results output by multiple matching expert modules for the same target sales data sample.
[0161] In one possible implementation, where the target loss function includes a first loss function and a second loss function, the training module 840 is specifically used for: Obtain multiple real sales tags corresponding to multiple target sales data samples; The first loss function is constructed based on multiple expert prediction results output by multiple preset expert modules and multiple real sales labels; A second loss function is constructed based on multiple expert prediction results output by multiple preset expert modules; The first loss function and the second loss function are weighted and combined to generate the target loss function.
[0162] In one possible implementation, the training module 840 is specifically used for: For each target sales data sample, multiple matching predicted sales information corresponding to the target sales data sample is obtained from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample. Multiple matching and predicted sales information are fused to generate target prediction results corresponding to target sales data samples, thereby obtaining multiple target prediction results corresponding to multiple target sales data samples; The first loss function is constructed based on multiple real sales labels and multiple target prediction results.
[0163] In one possible implementation, the training module 840 is specifically used for: Determine the expert weights of multiple matching expert modules corresponding to the target sales data sample; For each target sales data sample, multiple matching predicted sales information are weighted and fused based on multiple expert weights to obtain the target prediction result corresponding to the target sales data sample.
[0164] In one possible implementation, the training module 840 is specifically used for: For each target sales data sample, multiple matching predicted sales information corresponding to the target sales data sample is obtained from multiple expert prediction results output by multiple preset expert modules. The multiple matching predicted sales information are the results output by different matching expert modules for the target sales data sample. Based on multiple matching predicted sales information, inter-scenario prediction deviation information is generated. The inter-scenario prediction deviation information is used to indicate the degree of difference between any two matching predicted sales information for the same target sales data sample. A second loss function is constructed based on the prediction bias information between scenarios.
[0165] In one possible implementation, each sales data sample represents the sales record information of the corresponding target product within each preset sales period. The sales record information includes: the sales date of the target product and the actual sales volume of the target product within the preset sales period. The preset sales period is a daily sales period, and the preset time window contains multiple preset sales periods. The first determining module 820 is specifically used for: Based on each sales data sample, construct the sales statistics features of the target product within a preset time window. The sales statistics features include at least one of the following: sales summation feature, sales extreme value feature, sales trend feature, and sales dispersion feature. Based on the sales date of the target product represented by each sales data sample, a sales time feature corresponding to each sales data sample is constructed. The sales time feature includes at least one of the following: basic time feature, periodic time feature, and special time feature. The sales statistics features and sales time features are optimized to obtain the sample feature information corresponding to each sales data sample.
[0166] In one possible implementation, the preset initial model also includes a gated network module and a fusion network module; Training module 840 is specifically used for: Based on the objective loss function, the parameters of multiple preset expert modules, gated network modules and fusion network modules in the preset initial model are iteratively optimized until the trained preset initial model meets the preset convergence condition. Then, the trained preset initial model is determined as the sales prediction model.
[0167] In one possible implementation, multiple preset expert modules include a global expert module, which is used to process the sales data samples based on the sample feature information corresponding to each sales data sample in order to generate basic expert prediction results that do not depend on a specific scenario. The multiple preset expert modules also include at least one of the following: a promotional event scenario expert module, a new product scenario expert module, and an established product scenario expert module.
[0168] The division of modules in the above-described sales forecasting model training device is for illustrative purposes only. In other embodiments, the sales forecasting model training device can be divided into different modules as needed to complete all or part of the functions of the above-described sales forecasting model training device. The implementation of each module in the sales forecasting model training device provided in this application can be in the form of a computer program. This computer program can run on a terminal or server. The program modules constituted by this computer program can be stored in the memory of the terminal or server. When the computer program is executed by a processor, it implements all or part of the steps of the sales forecasting model training method described in this application.
[0169] The inventive concept based on the above-mentioned commodity sales forecasting method, such as Figure 9 As shown, this application also provides a product sales forecasting device 900 for implementing the product sales forecasting method described above. The product sales forecasting device 900 includes: The second acquisition module 910 is used to acquire historical sales data of the product to be predicted within a preset time window. The historical sales data includes sales record information of the product to be predicted in each preset sales cycle. The second determining module 920 is used to input historical sales data into the sales forecasting model to obtain the sales forecasting results of the sales forecasting model for the product to be predicted. The sales forecast model is generated by the training method of the sales forecast model provided in this application.
[0170] The division of modules in the above-described commodity sales forecasting device is for illustrative purposes only. In other embodiments, the commodity sales forecasting device can be divided into different modules as needed to complete all or part of the functions of the commodity sales forecasting device. The implementation of each module in the commodity sales forecasting device provided in this application can be in the form of a computer program. This computer program can run on a terminal or server. The program modules constituted by this computer program can be stored in the memory of the terminal or server. When the computer program is executed by a processor, it implements all or part of the steps of the commodity sales forecasting method described in this application.
[0171] This application also provides an electronic device, which can be a server, and its internal structure diagram can be as follows: Figure 10As shown, this electronic device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. The processor executes computer programs to implement a sales forecasting model training method or a product sales forecasting method.
[0172] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0173] In one possible implementation, a computer storage medium is provided that stores instructions, which, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the above embodiments. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium.
[0174] In one possible implementation, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0175] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer storage medium or transmitted through the computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).
[0176] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, sales record information involved in this specification was obtained with full authorization.
[0177] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0178] The above-described embodiments are merely preferred embodiments of this application and are not intended to limit the scope of this application. Any modifications and improvements made by those skilled in the art to the technical solutions of this application without departing from the spirit of this application shall fall within the protection scope defined by the claims.
[0179] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
Claims
1. A training method for a sales forecasting model, characterized in that, The method comprises the following steps: obtaining a preset initial model and a sales data sample set within a preset time window, the sales data sample set comprising a plurality of sales data samples corresponding to a plurality of target commodities; determining sample feature information corresponding to each sales data sample, the sample feature information comprising sales quantity statistical features and sales time features corresponding to the sales data sample; based on the sample feature information of each sales data sample, constructing a target mapping relationship between the plurality of sales data samples and a plurality of preset expert modules, wherein the plurality of preset expert modules are scenario-based expert sub-modules in the preset initial model; training the preset initial model according to the plurality of sales data samples and the target mapping relationship to obtain a sales quantity prediction model.
2. The method of claim 1, wherein, The target mapping relationship comprises a correspondence between each sales data sample and a plurality of matching expert modules; The method comprises the following steps: obtaining scenario feature information of each preset expert module; based on the sample feature information of each sales data sample and the scenario feature information of each preset expert module, determining each similarity information between the plurality of sales data samples and the plurality of preset expert modules; based on the similarity information, determining a plurality of matching expert modules corresponding to each sales data sample; based on each sales data sample and the plurality of matching expert modules corresponding thereto, constructing a target mapping relationship between the plurality of sales data samples and the plurality of preset expert modules.
3. The method of claim 1 or 2, wherein, The method comprises the following steps: determining a target sales data sample subset corresponding to each preset expert module according to the target mapping relationship, the target sales data sample subset comprising a plurality of target sales data samples matching the preset expert module; processing the corresponding target sales data sample subset through each preset expert module to obtain expert prediction results output by each preset expert module, the expert prediction results being used to represent a plurality of predicted sales quantity information corresponding to the plurality of target sales data samples; constructing a target loss function according to the plurality of expert prediction results output by the plurality of preset expert modules; training the preset initial model based on the target loss function to obtain a sales quantity prediction model.
4. The method of claim 3, wherein, The method comprises the following steps: for each preset expert module, obtaining a plurality of matching sales data samples corresponding to the preset expert module according to the target mapping relationship; based on an expert processing rule corresponding to the preset expert module, screening the plurality of matching sales data samples to obtain a plurality of target sales data samples matching the preset expert module, thereby generating a target sales data sample subset corresponding to the preset expert module.
5. The method of claim 3, wherein, The target loss function comprises a first loss function and / or a second loss function. The first loss function is used to represent an intra-scene prediction error between expert prediction results output by the respective preset expert module for the corresponding target sales data sample subset and a real sales label; and the second loss function is used to represent an inter-scene consistency deviation between multiple expert prediction results output by multiple matching expert modules for the same target sales data sample.
6. The method of claim 5, wherein, In a case where the target loss function comprises the first loss function and the second loss function, the constructing of the target loss function according to the multiple expert prediction results output by the multiple preset expert modules comprises: obtaining multiple real sales labels corresponding to the multiple target sales data samples; constructing a first loss function according to the multiple expert prediction results output by the multiple preset expert modules and the multiple real sales labels; constructing a second loss function according to the multiple expert prediction results output by the multiple preset expert modules; performing weighted combination on the first loss function and the second loss function to generate a target loss function.
7. The method of claim 6, wherein, The constructing of the first loss function according to the multiple expert prediction results output by the multiple preset expert modules and the multiple real sales labels comprises: for each target sales data sample, obtaining multiple matching predicted sales information corresponding to the target sales data sample from multiple expert prediction results output by the multiple preset expert modules, the multiple matching predicted sales information being results output by different matching expert modules for the target sales data sample; performing fusion on the multiple matching predicted sales information to generate a target prediction result corresponding to the target sales data sample, thereby obtaining multiple target prediction results corresponding to the multiple target sales data samples; constructing a first loss function according to the multiple real sales labels and the multiple target prediction results.
8. The method of claim 7, wherein, The performing of the fusion on the multiple matching predicted sales information to generate the target prediction result corresponding to the target sales data sample comprises: determining multiple expert weights of multiple matching expert modules corresponding to the target sales data sample; for each target sales data sample, performing weighted fusion on the multiple matching predicted sales information based on the multiple expert weights to obtain a target prediction result corresponding to the target sales data sample.
9. The method of claim 6, wherein, The constructing of the second loss function according to the multiple expert prediction results output by the multiple preset expert modules comprises: for each target sales data sample, obtaining multiple matching predicted sales information corresponding to the target sales data sample from multiple expert prediction results output by the multiple preset expert modules, the multiple matching predicted sales information being results output by different matching expert modules for the target sales data sample; generating inter-scene prediction deviation information based on the multiple matching predicted sales information, the inter-scene prediction deviation information being used to indicate a difference degree between any two matching predicted sales information for the same target sales data sample; constructing a second loss function according to the inter-scene prediction deviation information.
10. The method of claim 1 or 2, wherein, Each sales data sample represents sales record information of a corresponding target product in each preset sales period, and the sales record information includes a sales date of the target product and an actual sales volume of the target product in the preset sales period, the preset sales period is a daily sales period, and the preset time window includes a plurality of preset sales periods; The determination of the sample feature information corresponding to each sales data sample includes: Based on the sales data samples, a sales volume statistical feature of the target product in the preset time window is constructed, and the sales volume statistical feature includes at least one of a sales volume summation feature, a sales volume extreme value feature, a sales volume trend feature, and a sales volume dispersion feature; Based on the sales date of the target product represented by each sales data sample, a sales time feature corresponding to each sales data sample is constructed, and the sales time feature includes at least one of a basic time feature, a period time feature, and a special time feature; The sales volume statistical feature and the sales time feature are optimized to obtain sample feature information corresponding to each sales data sample.
11. The method of claim 3, wherein, The preset initial model further includes a gating network module and a fusion network module; The training of the preset initial model based on the target loss function to obtain a sales prediction model includes: Iterative optimization is performed on parameters of a plurality of preset expert modules, the gating network module, and the fusion network module in the preset initial model based on the target loss function, and when the trained preset initial model meets a preset convergence condition, the trained preset initial model is determined as the sales prediction model.
12. The method of claim 1 or 2, wherein, The plurality of preset expert modules include a global expert module, and the global expert module is configured to process the sales data samples based on the sample feature information corresponding to each sales data sample to generate a basic expert prediction result independent of a specific scenario; The plurality of preset expert modules further include at least one of a big promotion scenario expert module, a new product scenario expert module, and an old product scenario expert module.
13. A commodity sales forecasting method characterized by comprising: It includes: obtaining historical sales data of a to-be-predicted product in a preset time window, the historical sales data including sales record information of the to-be-predicted product in each preset sales period; inputting the historical sales data into a sales prediction model to obtain a product sales prediction result output by the sales prediction model for the to-be-predicted product; The sales prediction model is generated by the training method of the sales prediction model according to any one of claims 1 to 12. 14.A device for training a sales prediction model, comprising: It includes: A first obtaining module is configured to obtain a preset initial model and a sales data sample set in a preset time window, and the sales data sample set includes a plurality of sales data samples corresponding to a plurality of target products. A first determining module is configured to determine sample feature information corresponding to each sales data sample, and the sample feature information includes sales volume statistical features and sales time features corresponding to the sales data samples. The constructing module is configured to construct a target mapping relationship between the plurality of sales data samples and a plurality of preset expert modules based on sample feature information of the sales data samples, wherein the plurality of preset expert modules are scenario-based expert sub-models in the preset initial model. The training module is configured to train the preset initial model according to the plurality of sales data samples and the target mapping relationship to obtain a sales prediction model.
15. A commodity sales amount prediction device characterized by comprising: The method comprises the following steps: The second obtaining module is configured to obtain historical sales data of a to-be-predicted commodity within a preset time window, wherein the historical sales data comprises sales record information of the to-be-predicted commodity within each preset sales period. The second determining module is configured to input the historical sales data into a sales prediction model to obtain a commodity sales prediction result output by the sales prediction model for the to-be-predicted commodity. The sales prediction model is generated by the sales prediction model training method according to any one of claims 1 to 12.
16. An electronic device, comprising: The method comprises the following steps: A processor and a memory; The processor is connected with the memory; The memory is configured to store executable program codes; The processor runs a program corresponding to the executable program codes by reading the executable program codes stored in the memory, so as to execute the method according to any one of claims 1 to 13.
17. A computer storage medium, comprising, The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to execute the method steps according to any one of claims 1 to 13.
18. A computer program product comprising instructions, characterized in that, When the computer program product runs on a computer or a processor, the computer or the processor executes the method according to any one of claims 1 to 13.