Method for training supplementary quantity acquisition model, supplementary quantity acquisition method and device

By acquiring operational data from sample warehouses, using replenishment quantities to obtain model-estimated replenishment needs, and optimizing the model based on turnover and on-shelf data, the problem of inaccurate replenishment quantity calculation in e-commerce warehouses was solved, improving the accuracy of replenishment models and warehouse utilization.

CN121235596APending Publication Date: 2025-12-30ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511174153.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technology makes it difficult to accurately calculate the replenishment quantity of goods, which may lead to stockouts or low utilization rates in e-commerce warehouses.

Method used

By acquiring operational data from sample warehouses, the replenishment quantity acquisition model is used to estimate the required replenishment quantity, and reward values ​​are obtained based on the estimated turnover data and on-shelf data to optimize the replenishment quantity acquisition model.

Benefits of technology

It improves the accuracy and reliability of the replenishment volume acquisition model, enabling it to more effectively handle complex replenishment tasks and enhance warehouse utilization and merchandise flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121235596A_ABST
    Figure CN121235596A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for training a supplementary quantity acquisition model and a supplementary quantity acquisition method and device. And obtaining operation data of the sample object in the sample warehouse. And processing the operation data by using the supplementary quantity acquisition model to obtain the demand supplementary quantity of the sample object. And estimating turnover data of the sample objects and on-rack data of the sample objects after the sample warehouse is supplemented with the required supplement amount of the sample objects. And obtaining a reward value according to the pre-estimated turnover data and the pre-estimated on-shelf data. And optimizing the supplemental quantity acquisition model based on the reward value. In this way, the accuracy of the training target of the model can be improved, the reliability, the accuracy, the robustness and the generalization degree of the demand supplement amount output by the supplement amount obtaining model obtained through optimization based on the more accurate training target are higher, the accuracy of the replenishment amount obtained by the optimized supplement amount obtaining model can be improved, and the accuracy of the replenishment amount obtained by the optimized supplement amount obtaining model can be improved. And more complicated replenishment tasks can be more effectively coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and in particular to a method and apparatus for training a supplementary quantity acquisition model, and a supplementary quantity acquisition method and apparatus. Background Technology

[0002] Inventory management is an important technology in e-commerce. With the booming development of the e-commerce industry, inventory management has become a key factor affecting the operational efficiency and economic benefits of e-commerce.

[0003] Inventory management involves determining the replenishment quantity of goods. The accuracy of this replenishment quantity directly impacts whether e-commerce warehouses experience stockouts and their utilization rate, thus affecting profitability. Therefore, how to scientifically and accurately calculate the replenishment quantity of goods is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a method and apparatus for training a supplementary quantity acquisition model, and a supplementary quantity acquisition method and apparatus.

[0005] In a first aspect, this application discloses a method for training a model to obtain supplementary data, the method comprising:

[0006] Obtain operational data of sample objects in the sample repository;

[0007] The replenishment quantity acquisition model is used to process the operational data to obtain the required replenishment quantity of the sample object.

[0008] After the required replenishment quantity of the sample objects is estimated and added to the sample warehouse, the turnover data of the sample objects and the on-shelf data of the sample objects are obtained.

[0009] Rewards are awarded based on estimated turnover data and estimated inventory data.

[0010] The model for obtaining replenishment is optimized based on the reward value.

[0011] Secondly, this application discloses a method for obtaining a supplementary amount, the method comprising:

[0012] Obtain operational data for the object to be processed;

[0013] The operational data of the object to be processed is input into the replenishment acquisition model, so that the replenishment acquisition model processes the operational data to obtain the required replenishment amount of the object to be processed.

[0014] The replenishment acquisition model is obtained by optimizing the method shown in the first aspect.

[0015] Thirdly, this application discloses an apparatus for obtaining training supplementary data for a model, the apparatus comprising:

[0016] The first acquisition module is used to acquire operational data of sample objects in the sample repository;

[0017] The first processing module is used to process the operational data using the replenishment acquisition model to obtain the required replenishment amount of the sample object.

[0018] The estimation module is used to estimate the turnover data and the on-shelf data of the sample objects after the required replenishment quantity is added to the sample warehouse.

[0019] The second acquisition module is used to obtain reward values ​​based on the estimated turnover data and the estimated on-shelf data;

[0020] An optimization module is used to optimize the replenishment acquisition model based on the reward value.

[0021] Fourthly, this application discloses a supplementary quantity acquisition device, the device comprising:

[0022] The third acquisition module is used to acquire the operational data of the object to be processed;

[0023] The second processing module is used to input the operational data of the object to be processed into the replenishment acquisition model, so that the replenishment acquisition model processes the operational data to obtain the required replenishment amount of the object to be processed.

[0024] The replenishment acquisition model is obtained by optimizing the method shown in the first aspect.

[0025] The replenishment acquisition model is obtained by optimizing the device shown in the third aspect.

[0026] Fifthly, this application discloses an electronic device comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the methods shown in any of the foregoing aspects.

[0027] Sixthly, this application discloses a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods shown in any of the foregoing aspects.

[0028] In a seventh aspect, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the methods shown in any of the foregoing aspects.

[0029] This application has the following advantages:

[0030] In this application, operational data of sample objects in a sample warehouse are obtained. A replenishment acquisition model is used to process the operational data of the sample objects to obtain the required replenishment quantity for each sample object. After estimating the required replenishment quantity for sample objects in the sample warehouse, the turnover data and on-shelf data of the sample objects are obtained. A reward value is obtained based on the estimated turnover data and estimated on-shelf data. The replenishment acquisition model is then optimized based on the reward value.

[0031] Objects in a warehouse can be sold or leased. If the number of objects in the warehouse is insufficient, more objects need to be added to support subsequent sales or leases. In this scenario, the turnover data and on-shelf data of objects in the warehouse are often of concern. These indicators affect whether there is a shortage of objects in the warehouse and the utilization rate of the warehouse (for example, some objects that are stockpiled in the warehouse for a long time will lead to low warehouse utilization, while other objects will be sold or leased out of the warehouse quickly after entering the warehouse, which will lead to high warehouse utilization). Therefore, it is necessary to ensure that the turnover data of objects meets the expected requirements, and that the on-shelf data of sample objects meets the expected requirements.

[0032] Therefore, in this application, when obtaining the reward value, the estimated turnover data and estimated shelf data of the sample objects after replenishing the sample objects of the replenishment quantity obtained by the replenishment quantity model are considered. This allows for an indirect measurement of the reliability and accuracy of the replenishment quantity output by the replenishment quantity acquisition model. For example, it can indirectly measure whether the estimated turnover data and estimated shelf data of the sample objects meet the expected requirements after replenishing the sample objects of the replenishment quantity obtained by the replenishment quantity model in the sample warehouse. The reward value is then obtained based on the measurement results, and the replenishment quantity acquisition model is optimized based on the reward value. This improves the reliability and accuracy of the replenishment quantity output by the replenishment quantity acquisition model based on the strategy of "whether the estimated turnover data and estimated shelf data of the sample objects meet the expected requirements". This ensures that "after replenishing the objects of the replenishment quantity obtained by the replenishment quantity model in the warehouse, the estimated turnover data and estimated shelf data of the objects meet the expected requirements".

[0033] In summary, this application can improve the accuracy of the model's training objective. The replenishment quantity obtained by the model based on the more accurate training objective has higher reliability, accuracy, robustness, and generalization. Thus, it can improve the accuracy of the replenishment quantity obtained by the optimized replenishment quantity acquisition model, enabling it to cope with more complex replenishment tasks more effectively. Attached Figure Description

[0034] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of a system architecture according to this application.

[0036] Figure 2 This is a flowchart illustrating the steps of a method for obtaining training supplementary data for a model, as described in this application.

[0037] Figure 3 This is a schematic diagram of a method for obtaining training supplementary data for a model according to this application.

[0038] Figure 4 This is a flowchart illustrating the steps of one method for obtaining reward value according to this application.

[0039] Figure 5 This is a flowchart illustrating the steps of one method for obtaining reward value according to this application.

[0040] Figure 6 This is a flowchart illustrating the steps of one method for obtaining reward value according to this application.

[0041] Figure 7 This is a flowchart illustrating the steps of one method for obtaining reward value according to this application.

[0042] Figure 8 This is a flowchart of the steps for obtaining a supplementary quantity according to this application.

[0043] Figure 9 This is a structural block diagram of a device for obtaining training supplementary data for a model according to this application.

[0044] Figure 10 This is a structural block diagram of a supplementary quantity acquisition device according to this application.

[0045] Figure 11 This is a structural block diagram of a device according to this application. Detailed Implementation

[0046] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. This specification can be implemented in many other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific embodiments disclosed below. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort should fall within the scope of protection of this specification.

[0047] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments. Conversely, multiple steps described in this specification may be combined into a single step in other embodiments.

[0048] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0049] This application relates to neural network models and reinforcement learning in machine learning. By processing the operational data of sample objects through a neural network model, the required replenishment amount of the sample objects is obtained. After estimating the replenishment amount of sample objects in the sample warehouse, the turnover data and the on-shelf data of the sample objects are obtained. A reward value is obtained based on the estimated turnover data and the estimated on-shelf data. The parameters in the neural network model are optimized in reverse through reinforcement learning and the reward value.

[0050] See Figure 1 The diagram shows a system architecture diagram of this application.

[0051] The system includes: server, database, and user terminals.

[0052] The server and the user terminal can communicate directly or indirectly through wired or wireless communication methods, thus enabling data exchange between them via the network.

[0053] The server and the database can communicate directly or indirectly through wired or wireless communication methods, thus enabling data exchange between them over the network.

[0054] The database can store training data (such as operational data of sample objects), which is used to optimize the replenishment acquisition model.

[0055] The server can extract training data from the database to optimize the replenishment acquisition model. The optimized replenishment acquisition model is used to process the operational data of the sample objects to obtain the required replenishment amount of the objects, which can be used as a reference by the staff of the warehouse where the objects are located.

[0056] Alternatively, the server can encapsulate the replenishment acquisition model into a replenishment assistant application or an automatic replenishment system.

[0057] For example, a user terminal is a terminal used by warehouse staff. The staff may need to obtain the required replenishment quantity of a certain object in the warehouse. In this case, the staff can use the user terminal to send a request to the server to obtain the required replenishment quantity of that object.

[0058] When the server receives a request from the user terminal to obtain the replenishment amount for the object, the server can call the optimized replenishment amount acquisition model stored in the database to obtain the operational data of the object, input the operational data of the object into the replenishment amount acquisition model, so that the replenishment amount acquisition model processes the operational data of the object to obtain the replenishment amount of the object, and return the replenishment amount of the object to the user terminal so that the staff can obtain the replenishment amount of the object.

[0059] Alternatively, the server can send the optimized replenishment acquisition model to the user terminal and deploy the model on the user terminal. Then, when the user terminal needs to obtain the replenishment demand of a certain object in the warehouse, it can obtain the operational data of that object and input the operational data into the replenishment acquisition model deployed on the user terminal. This allows the replenishment acquisition model to process the operational data of the object and obtain the replenishment demand of that object, so that staff can obtain the replenishment demand of that object.

[0060] User terminals can include: mobile phones, tablets, laptops, desktop computers, PDAs, mobile internet devices (MIDs), and wearable devices, among other smart devices.

[0061] The server can be a standalone server or a group of servers.

[0062] The server can be a cloud server, also known as a cloud computing server or cloud host. It is a host product in the cloud computing service system, designed to solve the problems of high management difficulty and weak service scalability in traditional physical hosts and virtual private servers (VPS).

[0063] Before introducing the technical solution of this application, the technical categories that may be involved in the technical solution of this application will be explained.

[0064] DRL, Deep Reinforcement Learning, is an artificial intelligence technology that combines deep learning and reinforcement learning.

[0065] Replenishment decision: Based on current inventory, supplier response time, and other information, decide when and how much to purchase.

[0066] Mini-batch is a common training data processing method in machine learning and deep learning. It involves dividing the entire dataset into small batches (mini-batches) and then feeding them into the model batch by batch for training. This method is very popular in neural network training because it combines the advantages of both batch and stochastic gradient descent algorithms.

[0067] Ray is an open-source unified distributed computing framework that uses the `ray.remote` decorator, the Actor pattern, and a global scheduler to transform Python functions or classes into Tasks (stateless tasks) or Actors (stateful services) that can be executed in parallel across multiple machines. This supports various distributed training strategies such as data parallelism, model parallelism, and pipeline parallelism.

[0068] See Figure 2 This application illustrates a method for obtaining training supplementary data for a model, which may include:

[0069] In step S101, the operational data of the sample objects in the sample warehouse are obtained.

[0070] The sample repository of this application may be one or more.

[0071] A sample repository can contain one or more sample objects.

[0072] The sample objects in this application may include multiple sample objects, and the multiple sample objects may include sample objects of multiple kinds.

[0073] In this step, you can obtain the operational data of each sample object in multiple sample objects in a single sample repository, or, if there are multiple sample repositories, you can obtain the operational data of each sample object in multiple sample objects in each sample repository separately.

[0074] In summary, this step can obtain the operational data of each sample object in multiple sample objects, and then optimize the supplementary quantity acquisition model based on the operational data of each sample object in multiple sample objects (each round of optimization uses the operational data of one sample object or the operational data of two or more sample objects). The specific optimization method of one round of optimization of the supplementary quantity acquisition model using the operational data of the sample objects can be found in the description of steps S102 to S105, which will not be detailed here.

[0075] The sample objects include goods, such as goods for sale or lease in e-commerce warehouses (e.g., logistics warehouses used to store goods).

[0076] The sample includes goods in the warehouse that meet the expiration date requirements.

[0077] For example, the shelf life of fresh produce does not exceed a certain period (e.g., 85 days, 90 days, or 95 days, etc., which can be determined according to the actual situation, and this application does not limit it). Therefore, if the sample object is fresh produce, then the sample object is a product with a shelf life not exceeding a certain period.

[0078] Alternatively, for example, non-perishable goods have an expiration date that exceeds a certain time. Therefore, if the sample object is a non-perishable good, then the sample object is a good whose expiration date can exceed or equal to a certain time.

[0079] The aforementioned multiple types of sample objects can include multiple types of goods. In addition, multiple types of sample objects can also include goods from multiple different warehouses, etc.

[0080] The operational data of the sample object may include the sample object's historical operational data, such as the sample object's operational data within a historical time period. The duration of the historical time period can be determined according to the actual situation, and this application does not limit it. The end time of the historical time period is the time when step S101 is executed.

[0081] In one example, suppose we need to optimize the replenishment acquisition model using operational data of sample objects over a 90-day time interval in the historical process.

[0082] The day in which step S101 is executed can be the last day of the 90 days.

[0083] Each replenishment cycle is 7 days. Thus, sample objects need to be replenished in the warehouse on the 1st day, the 8th day, the 15th day, the 22nd day, the 29th day, and so on. Details will not be elaborated further.

[0084] Specifically, for the statement "sample objects need to be replenished in the warehouse on the 8th day", the required replenishment quantity of sample objects to be replenished in the warehouse on the 7th day (e.g., the night of the 7th day) needs to be determined. In this way, the operational data of the sample objects can be the operational data of the sample objects from the 1st to the 7th day.

[0085] Specifically, for the statement "sample objects need to be replenished in the warehouse on the 15th day", the required replenishment quantity of sample objects to be replenished in the warehouse on the 14th day (e.g., the night of the 14th day) needs to be determined. In this way, the operational data of the sample objects can be the operational data of the sample objects from the 1st to the 14th day.

[0086] Specifically, for the statement "sample objects need to be replenished in the warehouse on the 22nd day", the required replenishment quantity of sample objects to be replenished in the warehouse on the 22nd day needs to be determined on the 21st day (e.g., the night of the 21st day). In this way, the operational data of the sample objects can be the operational data of the sample objects from the 1st to the 21st day.

[0087] Specifically, for the statement "sample objects need to be replenished in the warehouse on the 29th day", the required replenishment quantity of sample objects to be replenished in the warehouse on the 29th day needs to be determined on the 28th day (e.g., the night of the 28th day). In this way, the operational data of the sample objects can be the operational data of the sample objects from the 1st to the 28th day.

[0088] ...and so on, without further details.

[0089] Additionally, each day that passes shifts the 90-day time interval forward by one day. For example, if the previous 90-day time interval was days 1-90, after one day it becomes days 2-91, and after another day it becomes days 3-92. Each shift of the 90-day time interval results in a new 90-day time interval, which can then be used to further optimize the replenishment acquisition model based on the new 90-day operational data of the sample objects.

[0090] Considering the universality of e-commerce scenarios and their relevance to product replenishment, the operational data of the sample objects can include multiple categories, such as the sample objects' basic data, the sample objects' historical sales data, the sample objects' future sales forecast data, the sample objects' purchase order (PO) data, and the sample objects' inventory data.

[0091] The basic data of the sample object is used to describe the basic attributes and characteristics of the sample object, and is the basis for the management and sales of the sample object.

[0092] The basic data of the sample object includes: the sample object's category, price, discount, packaging specifications, name, brand, code, detailed description, color, and function.

[0093] Historical sales data of the sample object is the sales record of the sample object over a period of time, which is used to analyze sales trends and assess market demand.

[0094] The historical sales data of the sample object includes: the sample object's sales volume for each day in the past, the sales volume of the sample object for each day in the past, the change rate of the sample object's sales volume for each day in the past, the relationship between the change rate of the sample object's sales volume for each day in the past and weekdays and holidays, and the sales volume of the sample object for each day in the past and the sales volume between promotional activities.

[0095] The future sales forecast data of the sample objects can help companies prepare inventory and make marketing plans in advance by estimating the future sales situation of the sample objects.

[0096] The future sales forecast data for the sample objects include: future sales forecasts based on promotional activities, future sales forecasts based on market trends, future sales forecasts based on competitor dynamics, and future sales forecasts based on new market expansion.

[0097] The purchase order data of the sample objects is information related to purchase orders, which is crucial for the sample objects' procurement management and coordination with suppliers.

[0098] The purchase order data for the sample objects includes: the name of the supplier, the address of the supplier, the contact information of the supplier, the rating of the supplier, the lead time of the purchase, the minimum purchase quantity, the purchase price, the purchase order number, the specifications of the purchased goods, the actual purchase quantity, and the delivery date.

[0099] The inventory data for the sample objects includes: the current inventory of the sample objects and the inventory in transit of the sample objects.

[0100] The operational data cited above is merely an example; in reality, it can include much more data to fully support the characteristics of various products and industries within the e-commerce scenario. This allows the replenishment acquisition model to optimize replenishment strategies based on sufficient information, resulting in stronger support, greater universality, and greater adaptability for various products and industries within the e-commerce scenario. Consequently, the optimized replenishment acquisition model can be a universal replenishment acquisition model applicable to multiple industries, multiple objectives, and multiple scenarios, exhibiting high generalization. A single replenishment acquisition model can be used to be compatible with multiple industries and adapt to different replenishment objectives, eliminating the need to optimize replenishment acquisition models separately for each industry, thereby reducing optimization and maintenance costs.

[0101] Using the operational data exemplified above can help the replenishment acquisition model learn appropriate parameters more quickly to output suitable replenishment amounts, thus reducing the search space and accelerating the convergence speed of the replenishment acquisition model.

[0102] If a sample object is out of stock on a particular day, meaning its inventory is 0, then the sample object cannot be sold on that day, and its actual sales volume is also 0.

[0103] However, in reality, if the sample object is in stock on that day, that is, the inventory is greater than 0, then according to past patterns, some of the sample object will still be sold on that day. In other words, the actual sales of the sample object will be greater than 0. Thus, by backtesting the history, we can determine that the sales of the sample object on that day should not be 0, but rather greater than 0. Specifically, we can use the average sales of the sample object over the previous days as the sales of the sample object on that day.

[0104] If the sales volume of a sample object on a particular day is significantly higher than the sales volume on other days—for example, if the sample object's daily sales volume is usually around 10 units, but the sales volume on a particular day is 10,000 units—then the sales volume on that day is often an anomaly, possibly due to statistical errors. Therefore, it is necessary to correct for the anomaly sales volume on that day. For example, the normal range for sales volume is μ ± 3σ. μ represents the expected sales volume of the sample object over multiple days. σ represents the standard deviation of the sales volume of the sample object over multiple days.

[0105] If the sales volume of the sample object on that day is not within the normal range of μ±3σ, then the sales volume of the sample object on that day will be converted to a value within the normal range of μ±3σ, for example, converted to "μ-3σ" or "μ+3σ".

[0106] Secondly, the operational data for any of the above categories can be normalized. For example, for any category of operational data, the maximum and minimum values ​​of that category's operational data for each sample object can be selected. For any sample object, the first difference between the operational data for that category and the minimum value can be calculated, the second difference between the maximum and minimum values ​​can be calculated, and the ratio between the first and second differences can be calculated. This ratio is then used as the normalized operational data for that category of the sample object for subsequent processing. The same applies to the operational data for each of the other categories.

[0107] In step S102, the replenishment quantity acquisition model is used to process the operational data of the sample objects to obtain the required replenishment quantity of the sample objects.

[0108] The input data for the replenishment quantity acquisition model is the operational data of the sample objects, and the output data is the replenishment quantity required by the sample objects. The replenishment quantity required by the sample objects is used to represent the sample objects that need to replenish the required quantity of goods in the sample warehouse.

[0109] Following the example in step S101, if the obtained operational data of the sample object is the operational data of the sample object from day 1 to day 7, then after the replenishment acquisition model processes the operational data of the sample object from day 1 to day 7, it obtains the required replenishment amount of the sample object that needs to be replenished in the warehouse on day 8 (e.g., in the early morning or morning of day 8).

[0110] If the operational data of the sample object obtained is the operational data of the sample object from day 1 to day 14, then after processing the operational data of the sample object from day 1 to day 14, the replenishment model obtains the required replenishment amount of the sample object that needs to be replenished in the warehouse on day 15 (e.g., in the early morning or morning of day 15).

[0111] If the operational data of the sample object obtained is the operational data of the sample object from day 1 to day 21, then after the replenishment acquisition model processes the operational data of the sample object from day 1 to day 21, it will obtain the required replenishment amount of the sample object that needs to be replenished in the warehouse on day 22 (e.g., early morning or morning of day 22).

[0112] If the operational data of the sample object obtained is the operational data of the sample object from day 1 to day 28, then after the replenishment acquisition model processes the operational data of the sample object from day 1 to day 28, it will obtain the required replenishment amount of the sample object that needs to be replenished in the warehouse on day 29 (e.g., in the early morning or morning of day 29).

[0113] ...and so on, without further details.

[0114] The supplementary data acquisition model calculates a first matrix of 5 rows and 1 column based on the operational data of the sample objects. The supplementary data acquisition model also has a second matrix, which is also a 5-row, 1-column matrix.

[0115] The transpose of the first matrix can be obtained, resulting in a 1x5 transpose matrix. The product of the 1x5 transpose matrix and the 5x1 second matrix is ​​then calculated to obtain a value, which represents the required supplementary quantity of the sample object.

[0116] The second matrix, consisting of 5 rows and 1 column, contains 5 elements: the average daily sales of the sample object in the most recent replenishment cycle * X * replenishment cycle, the average daily sales of the sample object in the most recent X replenishment cycles in history * Y, the average daily sales of the sample object in the most recent future replenishment cycle * X * replenishment cycle, the average daily sales of the sample object in the most recent X future replenishment cycles * Y, and a specific value.

[0117] Specific values ​​include 5, 10, 15 or 20, etc., which can be determined according to the circumstances, and this application does not limit them.

[0118] X is a positive integer greater than or equal to 2, and the replenishment cycle is a positive integer, for example, 7.

[0119] X * Replenishment Cycle = Y.

[0120] In step S103, after estimating the required replenishment quantity of sample objects in the sample warehouse, the turnover data and on-shelf data of the sample objects are obtained.

[0121] In this application, the estimated turnover data and estimated shelf data of the sample objects are data simulated based on the demand replenishment quantity. For example, a simulation tool can be used to simulate a scenario where the required replenishment quantity of sample objects is replenished in the sample warehouse. Combined with the actual sales volume and inventory of the sample objects, the turnover data and shelf data of the sample objects can be estimated.

[0122] The turnover data of the sample objects includes the turnover days of the sample objects, which describes how many days it takes for the sample objects in the warehouse to be sold out.

[0123] Turnover days = Sum of daily inventory quantities of sample objects in the sample warehouse over a period of time / Sum of sales volume of sample objects in the sample warehouse over a period of time. " / " is the division operator.

[0124] The on-shelf data for sample objects includes the on-shelf rate of sample objects, which describes the percentage of days that sample objects are present in the warehouse.

[0125] The shelf availability rate of the sample object = the number of days when the inventory of the sample object in the sample warehouse is greater than 0 / the total number of days. " / " is the division operator.

[0126] In one embodiment, the estimated turnover data of the sample object can be the historical and future turnover data of the sample object. The historical time period involved in the turnover data can be the same as the historical time period involved in the historical operation data of the sample object, and the future time period involved in the turnover data can be located after the historical time period involved in the historical operation data of the sample object and adjacent to the historical time period involved in the historical operation data of the sample object.

[0127] The estimated inventory data of the sample object can be the sample object's historical and future inventory data. The historical time period involved in the inventory data can be the same as the historical time period involved in the sample object's historical operating data. The future time period involved in the inventory data can be after the historical time period involved in the sample object's historical operating data and adjacent to the historical time period involved in the sample object's historical operating data.

[0128] The estimated turnover data for sample objects will utilize historical and future daily sales, inventory, and replenishment quantities (if a sample object is not sold on a certain day, its sales volume for that day is 0; if a sample object is not present in the sample warehouse on a certain day, meaning it is sold out, its inventory for that day is 0; if no sample object is replenished in the sample warehouse on a certain day, its replenishment demand for that day is 0). The estimated shelf-life data for sample objects will utilize historical and future daily sales, inventory, and replenishment quantities.

[0129] For example, following the example in step S102, if the replenishment acquisition model processes the operational data of the sample objects from day 1 to day 7 to obtain the required replenishment quantity of the sample objects that need to be replenished in the warehouse on day 8, then the estimated data is the turnover data of the sample objects from day 1 to day 14 after the required replenishment quantity of the sample objects is replenished in the sample warehouse on day 8, as well as the on-shelf data of the sample objects from day 1 to day 14.

[0130] The estimated sample item turnover data will use the sample item sales, inventory, and replenishment volume for each day from day 1 to day 14. The estimated sample item shelf data will use the sample item sales, inventory, and replenishment volume for each day from day 1 to day 14.

[0131] If the replenishment acquisition model processes the operational data of the sample objects from days 1 to 14 to obtain the required replenishment quantity of sample objects to be replenished in the warehouse on day 15, then the estimated data will be the turnover data and the on-shelf data of the sample objects from days 1 to 21 after the required replenishment quantity is replenished in the sample warehouse on day 15.

[0132] If the replenishment acquisition model processes the operational data of the sample objects from day 1 to day 21 to obtain the required replenishment quantity of sample objects to be replenished in the warehouse on day 22, then the estimated data will be the turnover data and the on-shelf data of the sample objects from day 1 to day 28 after the required replenishment quantity is replenished in the sample warehouse on day 22.

[0133] If the replenishment acquisition model processes the operational data of the sample objects from days 1 to 28 to obtain the required replenishment quantity of sample objects to be replenished in the warehouse on day 29, then the estimated data will be the turnover data of the sample objects from days 1 to 34 and the on-shelf data of the sample objects from days 1 to 35 after the required replenishment quantity is replenished in the sample warehouse on day 29.

[0134] In step S104, a reward value is obtained based on the estimated turnover data and the estimated on-shelf data.

[0135] The reward value can be a dynamic parameter. After processing the operational data of different sample objects using the replenishment quantity acquisition model, if the obtained demand replenishment quantities are not the same, the estimated turnover data and on-shelf data of the sample objects after estimating the replenishment demand in the sample warehouse will often be different. Thus, the estimated turnover data and estimated on-shelf data obtained based on different demand replenishment quantities will be different, and consequently, the reward values ​​will also be different. The performance of the replenishment quantity acquisition model is evaluated by using the estimated turnover data and estimated on-shelf data.

[0136] For details on how to obtain reward values ​​based on estimated turnover data and estimated inventory data, please refer to the embodiments shown later, which will not be described in detail here.

[0137] In step S105, the model for obtaining the replenishment amount is optimized based on the reward value.

[0138] The supplementary data acquisition model in this step is an unoptimized model, such as an initial model that has not yet started optimization, or an intermediate model obtained by performing at least one round of optimization (iteration) on the initial model (the parameters have not yet converged).

[0139] In a scenario where the replenishment acquisition model is optimized using the operational data of a sample object, the replenishment acquisition model is used to process the operational data to obtain the required replenishment amount of the sample object. After estimating the required replenishment amount of the sample object in the sample warehouse, the turnover data and the on-shelf data of the sample object are used. If the reward value obtained based on the estimated turnover data and the estimated on-shelf data is close to 0, it means that the required replenishment amount of the sample object obtained after processing the operational data using the replenishment acquisition model is more accurate (closer to the expected target).

[0140] Alternatively, in a scenario where the replenishment acquisition model is optimized using the operational data of a sample object, the replenishment acquisition model is used to process the operational data to obtain the required replenishment amount of the sample object. After estimating the required replenishment amount of the sample object in the sample warehouse, the turnover data and the on-shelf data of the sample object are used. If the reward value obtained based on the estimated turnover data and the estimated on-shelf data is greater than 0 and far from 0, it indicates that the required replenishment amount of the sample object obtained after processing the operational data using the replenishment acquisition model is less accurate (farther from the expected target).

[0141] When using the obtained reward value to update the replenishment amount to obtain the parameters in the model, unsupervised optimization methods or supervised optimization methods can be used.

[0142] For supervised optimization, the reward value obtained in the previous round can be used to calculate the loss value in the loss function (such as the cross-entropy loss function), and the replenishment amount can be adjusted according to the loss value to obtain the parameters in the model (the method of adjusting the replenishment amount to obtain the parameters in the model according to the loss value can refer to existing methods, and this application does not limit it). Then, the operational data of the next sample object is used to optimize the model with adjusted parameters for the next round of optimization.

[0143] In one example, methods for optimizing the replenishment acquisition model based on reward values ​​could include PPO (Proximal Policy Optimization) or GRPO (Group Relative Policy Optimization).

[0144] Thus, when there are multiple training data (operational data of multiple sample objects), multiple training data can be used to iteratively optimize the supplementary quantity acquisition model, continuously update the parameters in the supplementary quantity acquisition model, and the optimization can end when the parameters in the supplementary quantity acquisition model meet the convergence condition.

[0145] The replenishment acquisition model can be tested (e.g., using a large number of test objects of various types to test the replenishment acquisition model). Once the test results meet the requirements, the replenishment acquisition model can be deployed online and used to predict the required replenishment amount of objects in the repository.

[0146] The optimization objective for the replenishment acquisition model can be to minimize turnover data and maximize on-shelf data.

[0147] For example, given pre-set expected turnover data and expected inventory data, the optimization objective of the replenishment acquisition model could be: the turnover data estimated by the replenishment demand output from the replenishment acquisition model (the estimated turnover data gradually decreases in multiple rounds of optimization) should be close to or equal to the expected turnover data; and the inventory data estimated by the replenishment demand output from the replenishment acquisition model (the estimated inventory data gradually increases in multiple rounds of optimization) should be close to or equal to the expected inventory data. Thus, the convergence condition could be that the turnover data estimated by the replenishment demand output from the replenishment acquisition model is close to or equal to the expected turnover data, and the inventory data estimated by the replenishment demand output from the replenishment acquisition model is close to or equal to the expected inventory data.

[0148] Alternatively, the convergence condition can be: the number of optimizations to the supplementary quantity acquisition model is greater than or equal to a preset number, such as 500, 600, 1000 or 1500, which can be determined according to the actual situation. This application does not limit this.

[0149] In this application, operational data of sample objects in a sample warehouse are obtained. A replenishment acquisition model is used to process the operational data of the sample objects to obtain the required replenishment quantity for each sample object. After estimating the required replenishment quantity for sample objects in the sample warehouse, the turnover data and on-shelf data of the sample objects are obtained. A reward value is obtained based on the estimated turnover data and estimated on-shelf data. The replenishment acquisition model is then optimized based on the reward value.

[0150] Objects in a warehouse can be sold or leased. If the number of objects in the warehouse is insufficient, more objects need to be added to support subsequent sales or leases. In this scenario, the turnover data and on-shelf data of objects in the warehouse are often of concern. These indicators affect whether there is a shortage of objects in the warehouse and the utilization rate of the warehouse (for example, some objects are stockpiled in the warehouse for a long time, which will lead to low warehouse utilization, while other objects will be sold or leased out of the warehouse quickly after entering the warehouse, which will lead to high warehouse utilization). Therefore, it is necessary to ensure that the turnover data of objects meets the expected requirements, and that the on-shelf data of sample objects meets the expected requirements.

[0151] Therefore, in this application, when obtaining the reward value, the estimated turnover data and estimated shelf data of the sample objects after replenishing the sample objects of the replenishment quantity obtained by the replenishment quantity model are considered. This allows for an indirect measurement of the reliability and accuracy of the replenishment quantity output by the replenishment quantity acquisition model. For example, it can indirectly measure whether the estimated turnover data and estimated shelf data of the sample objects meet the expected requirements after replenishing the sample objects of the replenishment quantity obtained by the replenishment quantity model in the sample warehouse. The reward value is then obtained based on the measurement results, and the replenishment quantity acquisition model is optimized based on the reward value. This improves the reliability and accuracy of the replenishment quantity output by the replenishment quantity acquisition model based on the strategy of "whether the estimated turnover data and estimated shelf data of the sample objects meet the expected requirements". This ensures that "after replenishing the objects of the replenishment quantity obtained by the replenishment quantity model in the warehouse, the estimated turnover data and estimated shelf data of the objects meet the expected requirements".

[0152] In summary, this application can improve the accuracy of the model's training objective. The replenishment quantity obtained by the model based on the more accurate training objective has higher reliability, accuracy, robustness, and generalization. Thus, it can improve the accuracy of the replenishment quantity obtained by the optimized replenishment quantity acquisition model, enabling it to cope with more complex replenishment tasks more effectively.

[0153] This application calculates reward values ​​automatically and optimizes the replenishment quantity model based on the reward values, thereby reducing manual intervention and lowering labor costs.

[0154] The replenishment acquisition model in this application can be an end-to-end model. After it is put into use, the operational data of the object is input into it, and the replenishment acquisition model can process the operational data of the object to obtain the required replenishment amount of the object, which is the number of objects that need to be replenished in the warehouse. It has low computational complexity and is easy for warehouse staff to use.

[0155] End-to-end models can avoid errors and lower ceilings caused by various model assumptions, and are more likely to find suitable replenishment strategies.

[0156] Deep reinforcement learning excels at solving multi-period sequential decision optimization problems, performs well in inventory management, and can effectively improve indicators such as turnover data and on-shelf data.

[0157] in addition, Figure 3 A schematic diagram of an optimization method for a supplementary quantity acquisition model is shown.

[0158] Obtain operational data of sample objects from the sample repository and input it into the supplementary quantity acquisition model.

[0159] The replenishment acquisition model can process the operational data of the sample objects to obtain the required replenishment amount of the sample objects, and input the required replenishment amount into the forecasting tool.

[0160] The forecasting tool simulates a scenario where sample objects are replenished to meet demand in a sample warehouse. It forecasts the turnover and inventory data of these sample objects under this scenario and inputs this data into the reward evaluation center. Secondly, it can acquire various operational data, such as sales and inventory, within this scenario and update the operational data of the sample objects based on this data for later use.

[0161] The reward assessment center obtains reward values ​​based on the estimated turnover data and the estimated on-shelf data of the sample objects, and optimizes the replenishment acquisition model based on the reward values.

[0162] In one embodiment of this application, the optimization of the supplementary quantity acquisition model may involve multiple rounds of optimization. In each round of optimization, the operational data of a sample object can be used to optimize the supplementary quantity acquisition model. However, in each round of optimization, using only the operational data of a sample object to optimize the supplementary quantity acquisition model may sometimes lead to overfitting of the supplementary quantity acquisition model to that single sample object.

[0163] Therefore, in another embodiment of this application, the multi-round optimization is divided into two batches of optimization, the first batch of optimization including at least two rounds of optimization, and the second batch of optimization including at least two rounds of optimization.

[0164] In each round of optimization within the first batch of optimizations, operational data from a sample object can be used to optimize the supplementary quantity acquisition model.

[0165] After the first batch of optimization is completed, the second batch of optimization is carried out. In each round of optimization in the second batch, the model can be optimized based on the supplementary amount based on the minibatch.

[0166] For example, in each round of optimization in the second batch, operational data of multiple types of sample objects are used. These multiple types of sample objects include multiple types of products, such as obtaining operational data for hand cream, toothpaste, power banks, shampoo, and hair dryers, etc. Then, the supplementary quantity acquisition model is optimized in one round based on the operational data of multiple types of sample objects. That is, each round of optimization uses operational data of multiple types of sample objects.

[0167] Among them, the Ray mechanism can be used to optimize the supplementary quantity acquisition model in a round using operational data from multiple types of sample objects in the form of minibatch.

[0168] This embodiment can avoid overfitting of the supplementary quantity acquisition model and improve the accuracy of the supplementary quantity acquisition model.

[0169] Secondly, it can also speed up the optimization process. For example, using mini-batch can calculate gradients and update parameters in the supplementary acquisition model more quickly, which can improve the optimization efficiency of the supplementary acquisition model.

[0170] Using mini-batch updates for each parameter in the supplementary data acquisition model by averaging across multiple sample objects reduces instability in parameter updates, allowing the parameters in the supplementary data acquisition model to converge stably.

[0171] Furthermore, mini-batch processing can better utilize GPUs (Graphics Processing Units) or other parallel computing resources, and batch data processing can effectively match the hardware's parallel computing capabilities, improving optimization speed.

[0172] Based on mini-batch, operational data from various types of sample objects can be used together to optimize the replenishment quantity acquisition model. The optimized replenishment quantity acquisition model is applicable to various types of objects and has a relatively accurate replenishment quantity input capability for each type of object. It has strong generalization ability and can output more accurate replenishment quantities based on objects similar to the new type of object, even when facing new types of objects.

[0173] In one embodiment of this application, see [link to embodiment]. Figure 4 Step S104 includes:

[0174] In step S201, the expected turnover data of the sample object is obtained, and the turnover data deviation between the estimated turnover data and the expected turnover data is obtained.

[0175] Turnover data deviation is used to measure the degree to which the estimated turnover data matches or is close to the expected turnover data after replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the model. This indirectly measures the reliability and accuracy of the demand replenishment quantity output by the replenishment quantity acquisition model.

[0176] After replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the model, the more the estimated turnover data matches or is closer to the expected turnover data, the greater the reliability and accuracy of the demand replenishment quantity output by the replenishment quantity acquisition model.

[0177] Alternatively, after replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the model, the more the estimated turnover data is inconsistent with or close to the expected turnover data, the less reliable and accurate the demand replenishment quantity output by the replenishment quantity model will be.

[0178] The deviation between the estimated turnover data and the expected turnover data includes: the difference between the estimated turnover data and the expected turnover data, such as the difference between the estimated turnover data and the expected turnover data.

[0179] The expected turnover data for the sample objects includes: the expected turnover days for the sample objects.

[0180] The estimated turnover data for the sample objects includes: the estimated turnover days for the sample objects.

[0181] Thus, the turnover data deviation between the estimated turnover data and the expected turnover data includes: the difference between the estimated turnover days and the expected turnover days, for example, the difference between the estimated turnover days and the expected turnover days.

[0182] The expected turnover data of the sample objects are set in advance by the technicians according to the actual situation. For example, it can be 25, 26, 27, 30, 32 or 35, etc. The specific value can be determined according to the actual situation, and this application does not limit it.

[0183] Different types of objects have different expected turnover data. Therefore, by using their respective expected turnover data for different types of objects, the optimized replenishment quantity acquisition model can be adapted to each type of object.

[0184] In step S202, the expected inventory data of the sample object is obtained, and the inventory data deviation between the expected inventory data and the estimated inventory data is obtained.

[0185] The inventory data deviation is used to measure the degree to which the estimated inventory data matches or is close to the expected inventory data after replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the replenishment quantity acquisition model. This indirectly measures the reliability and accuracy of the demand replenishment quantity output by the replenishment quantity acquisition model.

[0186] After replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the model, the more the estimated inventory data matches or is closer to the expected inventory data, the greater the reliability and accuracy of the demand replenishment quantity output by the replenishment quantity acquisition model.

[0187] Alternatively, after replenishing the sample objects in the sample warehouse to obtain the demand replenishment quantity output by the model, the more the estimated inventory data is inconsistent with or close to the expected inventory data, the lower the reliability and accuracy of the demand replenishment quantity output by the replenishment quantity model.

[0188] The deviation between expected and estimated inventory data includes the difference between expected and estimated inventory data, such as the difference between expected and estimated inventory data.

[0189] The expected inventory data for the sample objects includes the expected inventory rate of the sample objects.

[0190] The estimated shelf availability data for the sample objects includes the estimated shelf availability rate of the sample objects.

[0191] Thus, the deviation between the expected inventory data and the estimated inventory data includes: the difference between the expected inventory rate and the estimated inventory rate, such as the difference between the expected inventory rate and the estimated inventory rate.

[0192] The expected data of the sample objects are set in advance by the technicians according to the actual situation. For example, it can be 0.9, 0.91, 0.92, 0.93, 0.95, 0.97 or 0.98, etc. The specific value can be determined according to the actual situation, and this application does not limit it.

[0193] The expected inventory data for different types of objects are not the same. Therefore, by using their respective expected inventory data for different types of objects, the optimized replenishment quantity acquisition model can be adapted to each type of object.

[0194] For steps S201 and S202, step S201 can be executed first and then step S202, or step S202 can be executed first and then step S201, or steps S201 and S202 can be executed in parallel.

[0195] In step S203, a reward value is obtained based on the turnover data deviation and the on-shelf data deviation.

[0196] In one embodiment of this application, the sum of the turnover data deviation and the on-shelf data deviation can be calculated and used as a reward value.

[0197] Alternatively, in another embodiment of this application, the reward value can be obtained by weighted summation of the turnover data deviation and the on-shelf data deviation.

[0198] For example, calculate the product of the turnover data deviation and its corresponding first weight, calculate the product of the on-shelf data deviation and its corresponding second weight, calculate the sum of these two products, and use it as the reward value.

[0199] The magnitudes of the first and second weights can be determined based on the actual situation. By using a weighted approach, the contribution of turnover data deviation and on-shelf data deviation to the reward value can be adjusted according to actual needs, increasing flexibility.

[0200] For example, if more importance is placed on turnover data deviation, the first weight corresponding to the turnover data deviation can be set to be larger, and greater than the second weight corresponding to the on-shelf data deviation. Alternatively, if more importance is placed on on-shelf data deviation, the second weight corresponding to the on-shelf data deviation can be set to be larger, and greater than the first weight corresponding to the turnover data deviation.

[0201] Alternatively, this step can also be described in the embodiments shown later, which will not be detailed here.

[0202] In one embodiment of this application, see [link to embodiment]. Figure 5 The process of obtaining the turnover data deviation between the estimated turnover data and the expected turnover data in step S201 includes:

[0203] In step S301, the first difference between the estimated turnover data and the expected turnover data is obtained.

[0204] One optimization objective of the replenishment acquisition model in this application is to minimize the turnover data. Thus, during the optimization process of the replenishment acquisition model, the estimated turnover data is gradually reduced within a range greater than the expected turnover data.

[0205] For example, if expected turnover data is set in advance, one optimization objective of the replenishment acquisition model is: the turnover data estimated based on the demand replenishment output by the replenishment acquisition model (the estimated turnover data gradually decreases from a value greater than the expected turnover data in multiple rounds of optimization) is close to or equal to the expected turnover data.

[0206] Therefore, in this step, the first difference between the estimated turnover data and the expected turnover data can be calculated, and then step S302 can be executed.

[0207] In step S302, the turnover data deviation is obtained based on the first difference.

[0208] In this application, the first difference can be defined as the turnover data deviation, or the value less than the first difference can be defined as the turnover data deviation.

[0209] For example, a comparison can be made between a first difference and a first preset value. If the first difference is less than the first preset value, the first preset value is determined as a turnover data deviation. Alternatively, if the first difference is greater than or equal to the first preset value, the first difference is determined as a turnover data deviation. That is, if the first preset value is larger, the first preset value is determined as a turnover data deviation; or, if the first difference is larger, the first difference is determined as a turnover data deviation.

[0210] The first preset value may include 0, 0.01, 0.02, 0.05 or 0.1, etc., and can be determined according to the actual situation. This application does not limit it in this regard.

[0211] For example, in this step, the turnover data deviation can be obtained based on the first difference using the following formula:

[0212] Turnover data deviation = max(first preset value, first difference).

[0213] Alternatively, the turnover data deviation = max(first preset value, estimated turnover data - expected turnover data)

[0214] In one example, the estimated turnover data is the estimated number of turnover days, the expected turnover data is the expected number of turnover days, and the first preset value is 0. Then, the turnover data deviation can be obtained as follows:

[0215] turnover_gap=max(0,turnover–turnover_target)

[0216] turnover_gap is the deviation of turnover data, turnover is the estimated turnover data, turnover_target is the expected turnover data, and max() is the calculation method to take the maximum value.

[0217] In one embodiment of this application, see [link to embodiment]. Figure 6 The process of obtaining the deviation between the expected and estimated inventory data in step S202 includes:

[0218] In step S401, a second difference between the expected inventory data and the estimated inventory data is obtained.

[0219] One optimization objective of the supplementary acquisition model in this application is to maximize the available data. Thus, during the optimization process of the supplementary acquisition model, the estimated available data gradually increases within a range that is less than the expected available data.

[0220] For example, if the expected inventory data is set in advance, one optimization objective of the replenishment acquisition model is: the inventory data estimated based on the demand replenishment output by the replenishment acquisition model (the estimated inventory data gradually increases from a value less than the expected inventory data in multiple rounds of optimization) is close to or equal to the expected inventory data.

[0221] Therefore, in this step, a second difference between the expected inventory data and the estimated inventory data can be calculated, and then step S402 can be executed.

[0222] In step S402, the on-shelf data deviation is obtained based on the second difference.

[0223] In this application, the second difference can be defined as the on-shelf data deviation, or a value less than the second difference can be defined as the on-shelf data deviation.

[0224] For example, the relationship between a second difference and a second preset value can be compared. If the second difference is less than the second preset value, the second preset value is determined as the on-shelf data deviation. Alternatively, if the second difference is greater than or equal to the second preset value, the second difference is determined as the on-shelf data deviation. That is, if the second preset value is larger, the second preset value is determined as the turnover on-shelf deviation; or, if the second difference is larger, the second difference is determined as the turnover data deviation.

[0225] The second preset value may include 0, 0.01, 0.02, 0.05 or 0.1, etc., and can be determined according to the actual situation. This application does not limit it in this regard.

[0226] For example, in this step, the deviation of the on-shelf data can be obtained based on the second difference using the following formula:

[0227] Data deviation = max(second preset value, second difference).

[0228] Alternatively, the deviation of inventory data = max(second preset value, expected inventory data - estimated inventory data)

[0229] In one example, the estimated inventory data is the estimated inventory rate, the expected inventory data is the expected inventory rate, and the first preset value is 0. Then, the deviation of the inventory data can be obtained as follows:

[0230] onshelf_gap=max(0, onshelf_target–onshelf)

[0231] onshelf_gap is the deviation of the on-shelf data, onshelf is the estimated on-shelf data, onshelf_target is the expected on-shelf data, and max() is the calculation method to take the maximum value.

[0232] In this application, the turnover data deviation is a single value, and the on-shelf data deviation is a single value. However, the order of magnitude of the turnover data deviation and the on-shelf data deviation are often different.

[0233] For example, suppose the turnover data deviation is the estimated turnover data minus the expected turnover data, and the expected turnover data is assumed to be 30. Since one of the optimization objectives of the replenishment acquisition model is that the turnover data estimated based on the demand replenishment output by the replenishment acquisition model (the estimated turnover data gradually decreases from a value greater than the expected turnover data in multiple rounds of optimization) is close to or equal to the expected turnover data, in the process of optimizing the replenishment acquisition model, the estimated turnover data is often greater than 30. Thus, the estimated turnover data minus the expected turnover data is also greater than 0, and usually greater than 1. For example, if the estimated turnover data is 80, 85, or 90, then the estimated turnover data minus the expected turnover data is 50, 55, or 60.

[0234] For example, suppose the deviation of the inventory data is the expected inventory data minus the estimated inventory data, and the expected inventory data is assumed to be 0.95. Since one of the optimization objectives of the replenishment acquisition model is that the inventory data estimated based on the demand replenishment output of the replenishment acquisition model (the estimated inventory data gradually increases from a value less than the expected inventory data in multiple rounds of optimization) is close to or equal to the expected inventory data, the estimated inventory data is often less than 0.95 during the optimization process of the replenishment acquisition model. Thus, the expected inventory data minus the estimated inventory data is usually greater than 1. For example, if the estimated inventory data is 0.9, 0.85, or 0.8, the expected inventory data minus the estimated inventory data is often 0.05, 0.1, or 0.15.

[0235] It can be seen that "estimated turnover data - expected turnover data = turnover data deviation (50, 55 or 60)" differs from "expected inventory data - estimated inventory data = inventory data deviation (0.05, 0.1 or 0.15)" by two orders of magnitude.

[0236] When the magnitudes of the turnover data deviation and the on-shelf data deviation are different, directly obtaining reward values ​​based on the turnover data deviation (50, 55, or 60) and the on-shelf data deviation (0.05, 0.1, or 0.15) will result in the turnover data deviation (50, 55, or 60) contributing far more than the on-shelf data deviation (0.05, 0.1, or 0.15) in reward calculation scenarios. The turnover data deviation (50, 55, or 60) severely reduces the contribution of the on-shelf data deviation (0.05, 0.1, or 0.15), and may even lead to a decrease in the contribution of the on-shelf data deviation. The contribution of the deviation (0.05, 0.1, or 0.15) is negligible, resulting in a small or even irrelevant correlation between the calculated reward value and the on-shelf data deviation (0.05, 0.1, or 0.15). This is equivalent to calculating the reward value without considering the on-shelf data deviation (0.05, 0.1, or 0.15), whereas this application requires consideration of the on-shelf data deviation (0.05, 0.1, or 0.15) when calculating the reward value. Therefore, the accuracy of the reward value obtained directly based on the turnover data deviation (50, 55, or 60) and the on-shelf data deviation (0.05, 0.1, or 0.15) is low.

[0237] Therefore, in order to improve the accuracy of the obtained reward value, in one embodiment of this application, it is necessary to normalize the magnitude of the turnover data deviation and the magnitude of the on-shelf data deviation to the same magnitude.

[0238] For example, see Figure 7 Step S203 includes:

[0239] In step S501, the normalized ratio of the order of magnitude between the turnover data deviation and the on-shelf data deviation is obtained.

[0240] In this application, turnover data can be measured in days (positive integers), while on-shelf data is measured as a percentage. In order to balance the deviation of turnover data and the deviation of on-shelf data, in this application, a 1-day deviation of turnover data can be treated as equivalent to a 1% deviation of on-shelf data. For example, the penalty for a 1-day increase in turnover data deviation is equal to the penalty for a 1% decrease in on-shelf data deviation.

[0241] In this application, when the order of magnitude of the turnover data deviation is higher than that of the in-shelf data deviation, a normalization ratio is obtained based on the order of magnitude difference between the order of magnitude of the turnover data deviation and the in-shelf data deviation. For example, if the order of magnitude difference is 1, the normalization ratio is 10 to the power of 1; if the order of magnitude difference is 2, the normalization ratio is 10 to the power of 2; if the order of magnitude difference is 3, the normalization ratio is 10 to the power of 3, and so on. If the order of magnitude difference is X, the normalization ratio is 10 to the power of X, where X is a positive integer.

[0242] For example, in the aforementioned example, the order of magnitude of the turnover data deviation (50, 55, or 60) is 10 to the power of 1, and the order of magnitude of the on-shelf data deviation (0.05, 0.1, or 0.15) is 10 to the power of -1. The order of magnitude difference between the two is 2. Thus, the normalization ratio can be 10 to the power of 2: 100.

[0243] In step S502, the turnover data deviation and the in-shelf data deviation are normalized according to the order of magnitude normalization ratio to obtain the normalized turnover data deviation and the normalized in-shelf data deviation.

[0244] In one embodiment, when the order of magnitude of the turnover data deviation is higher than that of the in-shelf data deviation, the turnover data deviation is directly used as the order of magnitude normalized turnover data deviation, and the product between the in-shelf data deviation and the normalization ratio is calculated to obtain the order of magnitude normalized in-shelf data deviation.

[0245] Alternatively, in another embodiment of this application, when the order of magnitude of the turnover data deviation is higher than that of the in-shelf data deviation, the ratio between the turnover data deviation and the normalization ratio is calculated and used as the order-of-magnitude normalized turnover data deviation, while the in-shelf data deviation is directly used as the order-of-magnitude normalized in-shelf data deviation.

[0246] In step S503, the reward value is obtained based on the normalized turnover data deviation and the normalized on-shelf data deviation.

[0247] In one embodiment of this application, the sum of the order-of-magnitude normalized turnover data deviation and the order-of-magnitude normalized on-shelf data deviation can be calculated and used as a reward value.

[0248] Alternatively, in another embodiment of this application, the order-of-magnitude normalized turnover data deviation and the order-of-magnitude normalized on-shelf data deviation can be weighted and summed to obtain the reward value. For example, the order-of-magnitude normalized turnover data deviation is calculated as the product of the order-of-magnitude normalized turnover data deviation and the first weight, and the order-of-magnitude normalized on-shelf data deviation is calculated as the product of the order-of-magnitude normalized on-shelf data deviation and the second weight. The sum of these two products is then used as the reward value.

[0249] In the foregoing embodiments, although the order of magnitude of the normalized turnover data deviation is the same as that of the normalized on-shelf data deviation, the numerical values ​​of the normalized turnover data deviation and the normalized on-shelf data deviation may differ significantly.

[0250] For example, there is a significant difference between the order-of-magnitude normalized turnover data deviation (50, 55, or 60) and the order-of-magnitude normalized on-shelf data deviation (5, 10, or 15).

[0251] Directly calculating reward values ​​based on the normalized turnover data deviation (50, 55, or 60) and the normalized on-shelf data deviation (5, 10, or 15) results in the normalized turnover data deviation (50, 55, or 60) contributing significantly more than the normalized on-shelf data deviation (5, 10, or 15) in reward calculation scenarios. The normalized turnover data deviation (50, 55, or 60) severely reduces the contribution of the normalized on-shelf data deviation (5, 10, or 15). This can even lead to the contribution of the normalized on-shelf data deviation (5, 10, or 15) being negligible. Consequently, the calculated reward value has little correlation with the normalized on-shelf data deviation (5, 10, or 15). However, this application requires full consideration of the normalized on-shelf data deviation (5, 10, or 15) when calculating the reward value. Therefore, the accuracy of the reward value obtained directly based on the normalized turnover data deviation (50, 55, or 60) and the normalized on-shelf data deviation (5, 10, or 15) is low.

[0252] Therefore, in order to improve the accuracy of the obtained reward value, in one embodiment of this application, it is necessary to normalize the value of the order-of-magnitude normalized turnover data deviation and the value of the order-of-magnitude normalized on-shelf data deviation to the same level.

[0253] Thus, in another embodiment of this application, the numerical normalization coefficient of the turnover data deviation can be obtained, and the turnover data deviation after order-of-magnitude normalization can be processed according to the numerical normalization coefficient of the turnover data deviation to obtain the numerically normalized turnover data deviation.

[0254] For example, the numerical normalization coefficient of turnover data deviation includes the expected turnover data of the sample object. In this way, the ratio between the order-of-magnitude normalized turnover data deviation and the expected turnover data of the sample object can be calculated and used as the numerically normalized turnover data deviation.

[0255] For example, the numerical normalization coefficient of turnover data deviation includes the uniform expected turnover data of each sample object. In this way, the ratio between the order-of-magnitude normalized turnover data deviation and the uniform expected turnover data can be calculated and used as the numerically normalized turnover data deviation.

[0256] Secondly, obtain the numerical normalization coefficient of the in-shelf data deviation, and perform numerical normalization processing on the in-shelf data deviation after order-of-magnitude normalization based on the numerical normalization coefficient of the in-shelf data deviation to obtain the numerically normalized in-shelf data deviation.

[0257] For example, the numerical normalization coefficient of the on-shelf data deviation includes the expected on-shelf data of the sample object. Thus, the ratio between the on-shelf data deviation normalized to the expected on-shelf data of the sample object can be calculated and used as the numerically normalized on-shelf data deviation.

[0258] For example, the numerical normalization coefficient of the on-shelf data deviation includes the uniform expected on-shelf data for each sample object. In this way, the ratio between the on-shelf data deviation normalized to the uniform expected on-shelf data can be calculated and used as the numerically normalized on-shelf data deviation.

[0259] Reward values ​​are derived based on the deviations in turnover data and on-shelf data after numerical normalization.

[0260] In one embodiment of this application, the sum of the turnover data deviation after numerical normalization and the on-shelf data deviation after numerical normalization can be calculated and used as a reward value.

[0261] Alternatively, in another embodiment of this application, the reward value can be obtained by weighted summation of the numerically normalized turnover data deviation and the numerically normalized on-shelf data deviation. For example, the product of the numerically normalized turnover data deviation and the first weight is calculated, the product of the numerically normalized on-shelf data deviation and the second weight is calculated, and the sum of these two products is used as the reward value.

[0262] Once the replenishment acquisition model has been optimized, it can be deployed online.

[0263] For example, see Figure 8 This application illustrates a method for obtaining a supplementary quantity, which may include:

[0264] In step S601, the operational data of the object to be processed is obtained.

[0265] For an explanation of the object to be processed, please refer to the explanation of the sample object in step S101, which will not be detailed here.

[0266] For an explanation of the operational data of the object to be processed, please refer to the explanation of the operational data of the sample object in step S101, which will not be detailed here.

[0267] In step S602, the operational data of the object to be processed is input into the replenishment acquisition model so that the replenishment acquisition model processes the operational data to obtain the required replenishment amount of the object to be processed.

[0268] The supplementary quantity acquisition model is obtained by optimizing any of the methods described above.

[0269] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.

[0270] Corresponding to the embodiments of the methods described above, this application also provides embodiments of the apparatus.

[0271] Reference Figure 9 The diagram shows a structural block diagram of an apparatus for obtaining training supplementation data for a model according to this application. The apparatus includes:

[0272] The first acquisition module 11 is used to acquire operational data of sample objects in the sample warehouse;

[0273] The first processing module 12 is used to process the operational data using the replenishment acquisition model to obtain the required replenishment amount of the sample object;

[0274] The estimation module 13 is used to estimate the turnover data and the on-shelf data of the sample objects after the required replenishment quantity of the sample objects is added to the sample warehouse.

[0275] The second acquisition module 14 is used to acquire reward values ​​based on the estimated turnover data and the estimated on-shelf data.

[0276] Optimization module 15 is used to optimize the supplementary amount acquisition model based on the reward value.

[0277] In one optional implementation, the second acquisition module includes:

[0278] The first acquisition unit is used to acquire the expected turnover data of the sample object; the second acquisition unit is used to acquire the turnover data deviation between the estimated turnover data and the expected turnover data.

[0279] The third acquisition unit is used to acquire the expected on-shelf data of the sample object; the fourth acquisition unit is used to acquire the on-shelf data deviation between the expected on-shelf data and the estimated on-shelf data.

[0280] The fifth acquisition unit is used to acquire the reward value based on the turnover data deviation and the on-shelf data deviation.

[0281] In one optional implementation, the second acquisition unit includes:

[0282] The first acquisition subunit is used to acquire the first difference between the estimated turnover data and the expected turnover data;

[0283] The second acquisition subunit is used to acquire the turnover data deviation based on the first difference.

[0284] In one optional implementation, the second acquisition subunit is specifically used to: determine the first difference as the turnover data deviation, or determine a value less than the first difference as the turnover data deviation.

[0285] In one optional implementation, the second acquisition subunit is specifically used to: compare the magnitude relationship between the first difference and the first preset value; in response to the first difference being less than the first preset value, determine the first preset value as the turnover data deviation; or, in response to the first difference being greater than or equal to the first preset value, determine the first difference as the turnover data deviation.

[0286] In one optional implementation, the fourth acquisition unit includes:

[0287] The third acquisition subunit is used to acquire the second difference between the expected on-shelf data and the estimated on-shelf data;

[0288] The fourth acquisition subunit is used to acquire the on-shelf data deviation based on the second difference.

[0289] In one optional implementation, the fourth acquisition subunit is specifically used to: determine the second difference as the on-shelf data deviation, or determine a value less than the second difference as the on-shelf data deviation.

[0290] In one optional implementation, the fourth acquisition subunit is specifically used to: compare the magnitude relationship between the second difference and the second preset value; in response to the second difference being less than the second preset value, determine the second preset value as the on-shelf data deviation; or, in response to the second difference being greater than or equal to the second preset value, determine the second difference as the on-shelf data deviation.

[0291] In one optional implementation, the fifth acquisition unit includes:

[0292] The weighted summation subunit is used to sum the turnover data deviation and the on-shelf data deviation in a weighted manner to obtain the reward value.

[0293] In one alternative implementation, the magnitude of the turnover data deviation is different from the magnitude of the on-shelf data deviation;

[0294] The fifth acquisition unit includes:

[0295] The fifth acquisition subunit is used to acquire the normalized ratio of the order of magnitude between the turnover data deviation and the on-shelf data deviation;

[0296] The normalization subunit is used to normalize the turnover data deviation and the on-shelf data deviation according to the order-of-magnitude normalization ratio, so as to obtain the order-of-magnitude normalized turnover data deviation and the order-of-magnitude normalized on-shelf data deviation.

[0297] The sixth acquisition subunit is used to acquire the reward value based on the order-of-magnitude normalized turnover data deviation and the order-of-magnitude normalized on-shelf data deviation.

[0298] In an optional implementation, the sixth acquisition subunit is specifically used for:

[0299] Obtain the numerical normalization coefficient of the turnover data deviation, and perform numerical normalization processing on the turnover data deviation after order of magnitude normalization based on the numerical normalization coefficient of the turnover data deviation to obtain the numerically normalized turnover data deviation.

[0300] Obtain the numerical normalization coefficient of the in-shelf data deviation, and perform numerical normalization processing on the in-shelf data deviation after order of magnitude normalization based on the numerical normalization coefficient of the in-shelf data deviation to obtain the numerically normalized in-shelf data deviation.

[0301] The reward value is obtained based on the deviation of the turnover data after numerical normalization and the deviation of the on-shelf data after numerical normalization.

[0302] Reference Figure 10 The diagram shows a structural block diagram of a supplementary quantity acquisition device according to this application, the device comprising:

[0303] The third acquisition module 21 is used to acquire the operational data of the object to be processed;

[0304] The second processing module 22 is used to input the operational data of the object to be processed into the replenishment acquisition model, so that the replenishment acquisition model processes the operational data to obtain the required replenishment amount of the object to be processed.

[0305] The supplementary quantity acquisition model is obtained based on any of the aforementioned optimization methods.

[0306] Alternatively, the supplementary quantity acquisition model may be obtained based on optimization using any of the aforementioned optimization devices.

[0307] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0308] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0309] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.

[0310] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more methods as described in the above embodiments. In this application, the electronic device includes a server, a gateway, sub-devices, etc., and the sub-devices are devices such as Internet of Things (IoT) devices.

[0311] Embodiments of this disclosure can be implemented as an apparatus with any suitable hardware, firmware, software, or any combination thereof, configured as desired. This apparatus may include electronic devices such as servers (clusters) and terminal devices such as IoT devices.

[0312] Figure 11 An exemplary apparatus 1300 is schematically shown that can be used to implement the various embodiments of this application.

[0313] In one embodiment, Figure 11 An exemplary device 1300 is shown, which includes one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0314] Processor 1302 may include one or more single-core or multi-core processors, and processor 1302 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1300 can function as a server device such as a gateway in the embodiments of this application.

[0315] In some embodiments, apparatus 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 that are combined with the one or more computer-readable media and configured to execute the instructions 1314 to implement the module and thus perform the actions in this disclosure.

[0316] In one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1302 and / or any suitable device or component communicating with the control module 1304.

[0317] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0318] Memory 1306 may be used, for example, to load and store data and / or instructions 1314 for device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1306 may include double data rate quad synchronous dynamic random access memory (DDR4 SDRAM).

[0319] In one embodiment, the control module 1304 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1308 and (one or more) input / output devices 1310.

[0320] For example, NVM / storage device 1308 may be used to store data and / or instructions 1314. NVM / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0321] NVM / storage device 1308 may include storage resources that are physically part of a device on which device 1300 is mounted, or that can be accessed by the device without needing to be part of the device. For example, NVM / storage device 1308 may be accessed via a network via one or more input / output devices 1310.

[0322] One or more input / output devices 1310 may provide an interface for device 1300 to communicate with any other suitable device. Input / output devices 1310 may include communication components, pinyin components, sensor components, etc. Network interface 1312 may provide an interface for device 1300 to communicate via one or more networks. Device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.

[0323] In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1304. In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers of the control module 1304 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die to form a system-on-a-chip (SoC).

[0324] In various embodiments, device 1300 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1300 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0325] This application provides an electronic device, including: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform one or more methods as described in this application.

[0326] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0327] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable information processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable information processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0328] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable information processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0329] These computer program instructions can also be loaded onto a computer or other programmable information processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0330] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0331] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.

[0332] The various methods and apparatuses provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for training a supplementary quantity acquisition model, characterized in that, The method comprises: acquiring operation data of a sample object in a sample warehouse; processing the operation data using a replenishment quantity acquisition model to obtain a demand replenishment quantity of the sample object; estimating turnover data of the sample object and on-shelf data of the sample object after replenishing the demand replenishment quantity of the sample object in the sample warehouse; acquiring a reward value according to the estimated turnover data and the estimated on-shelf data; optimizing the replenishment quantity acquisition model based on the reward value.

2. The method of claim 1, wherein, The acquiring of the reward value according to the estimated turnover data and the estimated on-shelf data comprises: acquiring expected turnover data of the sample object, and acquiring a turnover data deviation between the estimated turnover data and the expected turnover data; acquiring expected on-shelf data of the sample object, and acquiring an on-shelf data deviation between the expected on-shelf data and the estimated on-shelf data; acquiring the reward value according to the turnover data deviation and the on-shelf data deviation.

3. The method of claim 2, wherein, The acquiring of the turnover data deviation between the estimated turnover data and the expected turnover data comprises: acquiring a first difference value between the estimated turnover data and the expected turnover data; acquiring the turnover data deviation according to the first difference value.

4. The method of claim 3, wherein, The acquiring of the turnover data deviation according to the first difference value comprises: determining the first difference value as the turnover data deviation, or determining a value smaller than the first difference value as the turnover data deviation.

5. The method of claim 4, wherein, The determining of the first difference value as the turnover data deviation, or the determining of a value smaller than the first difference value as the turnover data deviation comprises: comparing a size relationship between the first difference value and a first preset value; determining the first preset value as the turnover data deviation in response to the first difference value being smaller than the first preset value; or determining the first difference value as the turnover data deviation in response to the first difference value being greater than or equal to the first preset value. The acquiring of the on-shelf data deviation between the estimated on-shelf data and the expected on-shelf data comprises:

6. The method of claim 2, wherein, acquiring a second difference value between the expected on-shelf data and the estimated on-shelf data; acquiring the on-shelf data deviation according to the second difference value. The acquiring of the on-shelf data deviation according to the second difference value comprises:

7. The method of claim 6, wherein, determining the second difference value as the on-shelf data deviation, or determining a value smaller than the second difference value as the on-shelf data deviation. The determining of the second difference value as the on-shelf data deviation, or the determining of a value smaller than the second difference value as the on-shelf data deviation comprises:

8. The method of claim 7, wherein, comparing a size relationship between the second difference value and a second preset value; determining the second preset value as the on-shelf data deviation in response to the second difference value being smaller than the second preset value; or determining the second difference value as the on-shelf data deviation in response to the second difference value being greater than or equal to the second preset value. The acquiring of the reward value according to the turnover data deviation and the on-shelf data deviation comprises: weighting and summing the turnover data deviation and the on-shelf data deviation to obtain the reward value.

9. The method of claim 2, wherein, ​ ​ 10. The method of claim 2, wherein, The order of magnitude of the value of the turnover data bias is different from the order of magnitude of the value of the on-shelf data bias; The reward value is obtained according to the turnover data bias and the on-shelf data bias, and the method comprises: An order of magnitude normalization ratio between the turnover data bias and the on-shelf data bias is obtained; The turnover data bias and the on-shelf data bias are subjected to order of magnitude normalization processing according to the order of magnitude normalization ratio, to obtain order of magnitude normalized turnover data bias and order of magnitude normalized on-shelf data bias; The reward value is obtained according to the order of magnitude normalized turnover data bias and the order of magnitude normalized on-shelf data bias.

11. The method of claim 10, wherein, The reward value is obtained according to the order of magnitude normalized turnover data bias and the order of magnitude normalized on-shelf data bias, and the method comprises: A value normalization coefficient of the turnover data bias is obtained, and the order of magnitude normalized turnover data bias is subjected to value normalization processing according to the value normalization coefficient of the turnover data bias, to obtain value normalized turnover data bias; A value normalization coefficient of the on-shelf data bias is obtained, and the order of magnitude normalized on-shelf data bias is subjected to value normalization processing according to the value normalization coefficient of the on-shelf data bias, to obtain value normalized on-shelf data bias; The reward value is obtained according to the value normalized turnover data bias and the value normalized on-shelf data bias.

12. A method of obtaining a replenishment quantity, characterized by The method comprises: Obtaining operation data of a to-be-processed object; Inputting the operation data of the to-be-processed object into a replenishment quantity obtaining model, so that the replenishment quantity obtaining model processes the operation data to obtain a demand replenishment quantity of the to-be-processed object; The replenishment quantity obtaining model is optimized based on the method in any one of claims 1-11.

13. An apparatus for obtaining training supplementation data for a model, characterized in that, The device comprises: A first obtaining module configured to obtain operation data of a sample object in a sample warehouse; A first processing module configured to process the operation data using a replenishment quantity obtaining model to obtain a demand replenishment quantity of the sample object; An estimation module configured to estimate turnover data of the sample object and on-shelf data of the sample object after replenishing the sample object with the demand replenishment quantity in the sample warehouse; A second obtaining module configured to obtain a reward value according to the estimated turnover data and the estimated on-shelf data; An optimization module configured to optimize the replenishment quantity obtaining model based on the reward value.

14. A replenishment quantity acquisition device characterized by comprising: The device comprises: A third obtaining module configured to obtain operation data of a to-be-processed object; A second processing module configured to input the operation data of the to-be-processed object into a replenishment quantity obtaining model, so that the replenishment quantity obtaining model processes the operation data to obtain a demand replenishment quantity of the to-be-processed object; The replenishment quantity obtaining model is optimized based on the method in any one of claims 1-11. Alternatively, the replenishment quantity obtaining model is optimized based on the device in claim 13.

15. An electronic device, comprising: A computer program product comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method in any one of claims 1-12 when executing the program. A computer program product comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method in any one of claims 1-12 when executing the program.

16. A computer-readable storage medium, characterized in that, A computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement the method in any one of claims 1 to 12.

17. A computer program product, characterised in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device is enabled to perform the method in any one of claims 1 to 12.