Warehouse input and output quantity prediction method and system based on integrated machine learning

Through the integrated machine learning method, multiple subsets are used to train independent basic learners to generate multiple prediction models, and the model with the decision coefficient closest to 1 is selected as the optimal prediction model, which solves the generalization and volatility problems in warehouse inlet and exit prediction, and improves the accuracy and robustness of prediction.

CN120338660APending Publication Date: 2025-07-18UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415866.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prediction effect of the prior art in warehouse entry and exit forecast needs to be optimized. The data materials are strict and have weak generalization, making it difficult to cope with high volatility and uncertainty.

Method used

Using an integrated machine learning method, by acquiring multiple subsets and training independent basic learners on each subset, multiple basic prediction models are generated, and the model with the decision coefficient closest to 1 is selected as the optimal prediction model, reducing the high variance problem and improving the robustness of the prediction.

Benefits of technology

It effectively reduces the high variance problem of a single model, improves the accuracy and robustness of predictions, and is suitable for scenarios with high uncertainty and high volatility in warehouse inlets and exits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338660A_ABST
    Figure CN120338660A_ABST
Patent Text Reader

Abstract

The invention discloses a warehouse input and output prediction method based on integrated machine learning. The method comprises the following steps: acquiring original sample data from a warehouse end; obtaining a plurality of subsets of the original sample data, and dividing each subset into training data and test data; training a pre-configured and independent basic learning device by using each group of training data to obtain a plurality of basic prediction models; performing basic prediction on each group of test data by utilizing each basic prediction model, and outputting a plurality of basic prediction values of input and output; calculating a decision coefficient between the basic prediction value of the input and output quantity and an actual value in the test data; selecting the basic prediction model with the decision coefficient closest to 1 as an optimal prediction model; and predicting the warehouse input and output based on the optimal prediction model. The invention also provides a prediction system. According to the method, the problem of high variance possibly existing in a single model is effectively reduced, and the method is particularly suitable for the characteristics of high uncertainty and large volatility of warehouse input and output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of logistics, and in particular, to a method and system for predicting the inbound and outbound volume of a warehouse integrated with machine learning. Background Art

[0002] Predicting the inbound and outbound volume of a warehouse refers to analyzing and predicting the inbound and outbound volume of goods or materials in the warehouse in order to better manage inventory and optimize the supply chain. Through accurate prediction of the inbound and outbound volume, the supply-demand relationship can be better balanced, the risk of inventory backlog or out-of-stock can be reduced, and the operational efficiency and customer satisfaction can be improved.

[0003] In order to achieve more accurate prediction, the prior art has begun to use artificial intelligence algorithms to model a large amount of complex historical sales data and combine external factors, such as weather conditions or holiday effects, etc. to improve the performance of the model.

[0004] However, due to current technical constraints, the prediction effect of the methods adopted needs to be further optimized, the data materials required for prediction are relatively strict, special devices are needed for storage, and the generalizability is weak. Summary of the Invention

[0005] In view of the above problems, a first aspect of the present application provides a method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning, including: obtaining original sample data from the warehouse side; the original sample data at least includes: the total volume of goods inbound and outbound in a first set unit period, the total quantity of goods inbound and outbound in a first set unit period, the outbound time of the goods or the inbound time of the goods; obtaining multiple subsets of the original sample data, and dividing each subset into training data and test data; using each set of training data to train a pre-configured and independent base learner to obtain multiple base prediction models; using each of the base prediction models to perform a base prediction on each set of the test data and outputting multiple basic predicted values of the inbound and outbound volume; wherein, the basic predicted values of the inbound and outbound volume include the predicted total volume of the inbound and outbound goods and / or the predicted total quantity of the inbound and outbound goods; calculating the coefficient of determination between the basic predicted values of the inbound and outbound volume and the actual values in the test data; selecting the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model; predicting the inbound and outbound volume of the warehouse based on the optimal prediction model; the inbound and outbound volume of the warehouse includes the total volume of the inbound and outbound goods and / or the total quantity of the inbound and outbound goods.

[0006] The second aspect of the present application provides a warehouse inbound and outbound volume prediction system based on integrated machine learning, including: a processing device, which includes: a first acquisition module configured to acquire original sample data from the warehouse side; the original sample data at least includes: the total volume of goods inbound and outbound in a first set unit period, the total quantity of goods inbound and outbound in a first set unit period, the outbound time of the goods or the inbound time of the goods; a division module configured to acquire multiple subsets of the original sample data and divide each subset into training data and test data; a training module configured to train a pre-configured and independent base learner using each set of training data to obtain multiple base prediction models; a base prediction module configured to perform base prediction on each set of the test data using each of the base prediction models and output multiple inbound and outbound volume base prediction values; wherein, the inbound and outbound volume base prediction values include the predicted total volume of inbound and outbound goods and / or the predicted total quantity of inbound and outbound goods; a verification module configured to calculate the coefficient of determination between the inbound and outbound volume base prediction values and the actual values in the test data; a selection module configured to select the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model; an actual prediction module configured to predict the warehouse inbound and outbound volume based on the optimal prediction model; the warehouse inbound and outbound volume includes the total volume of inbound and outbound goods and / or the total quantity of inbound and outbound goods.

[0007] The present application provides a method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning with a Bagging structure. By acquiring multiple subsets, training an independent base learner on each subset to obtain multiple base prediction models, and comparing the results of the base prediction models to obtain the optimal prediction model and further obtain the final output, it can effectively reduce the high variance problem that may exist in a single model, that is, overfitting the noise data in a specific training set, and is particularly suitable for the problem of high uncertainty and large volatility in the inbound and outbound volume of a warehouse. Training based on multiple subsets enables the prediction process to have the ability to resist outliers and noise. Even if some subsets contain inaccurate data points, the impact they cause will be offset by the good performance of other subsets.

[0008] After reading the specific embodiments of the present invention in conjunction with the accompanying drawings, other features and advantages of the present invention will become clearer. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1Flow chart of the warehouse inbound and outbound volume prediction method based on integrated machine learning provided by the present invention;

[0011] Figure 2 Flow chart of the warehouse inbound and outbound volume prediction method based on integrated machine learning provided by the present invention;

[0012] Figure 3 Schematic block diagram of the structure of the warehouse inbound and outbound volume prediction system based on integrated machine learning provided by the present invention;

[0013] Figure 4 Schematic block diagram of the structure of the warehouse inbound and outbound volume prediction system based on integrated machine learning provided by the present invention. Detailed implementation manners

[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0015] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0016] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "electrically connected" should be understood in a broad sense. For example, it can be a fixed electrical connection, a detachable electrical connection, or an integral electrical connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. In the description of the implementation manners, specific features, structures, materials, or characteristics can be combined in a suitable manner in any embodiment or example.

[0017] The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0018] In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.

[0019] The first aspect of the present application provides a method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning, including multiple steps as Figure 1 shown.

[0020] Step S101: Obtain the original sample data from the warehouse side.

[0021] In some embodiments of the present application, the original sample data at least includes: the total volume of the goods inbound and outbound in the first set unit period, the total number of the goods inbound and outbound in the first set unit period, the outbound time of the goods or the inbound time of the goods.

[0022] In some embodiments of the present application, the first set unit period may be a period in units of "days".

[0023] The original sample data can be directly obtained from the daily operation data of the warehouse or calculated based on the daily operation data of the warehouse. Exemplarily, the data of the following fields can be directly obtained from the daily operation data of the warehouse: "the outbound time of the goods", "the inbound time of the goods", "order number", "number of orders", "name of the originating warehouse", "name of the receiving warehouse", "total number of the goods in the order", "total volume of the goods in the order". Further, the total volume of the goods outbound daily, the total volume of the goods inbound daily, the total number of the goods inbound daily, and the total number of the goods outbound daily can be calculated, where the total volume of the goods outbound daily and the total volume of the goods inbound daily are the total volume of the goods inbound and outbound in the first set unit period, and the total number of the goods outbound daily and the total number of the goods inbound daily are the total number of the goods inbound and outbound in the first set unit period.

[0024] The length of the first set unit period can be set according to the actual needs according to the storage scale of the warehouse.

[0025] In some other embodiments of the present application, the original sample data may further include "order number", "number of orders", "name of the originating warehouse", "name of the receiving warehouse", "total number of the goods in the order", "total volume of the goods in the order" as input features.

[0026] The preprocessing process of the original sample data can adopt mature preprocessing methods in the prior art, such as cleaning, filtering, deleting missing values, etc., which will not be elaborated here.

[0027] Step S102: Obtain multiple subsets of the original sample data, and divide each subset into training data and test data.

[0028] In some embodiments of the present application, a number of subsets are drawn from the original sample data by means of random sampling with replacement. A preferred drawing method will be introduced in detail below.

[0029] Step S103: Train a pre-configured and independent base learner using each set of training data to obtain a plurality of base prediction models.

[0030] Step S104: Use each base prediction model to perform a base prediction on each set of test data and output the base prediction value of the inbound and outbound volume for the first set unit period.

[0031] In some embodiments of the present application, the base prediction value of the inbound and outbound volume includes the predicted total volume of the inbound and outbound goods.

[0032] In some embodiments of the present application, the base prediction value of the inbound and outbound volume includes the predicted total quantity of the inbound and outbound goods.

[0033] In some embodiments of the present application, the base prediction value of the inbound and outbound volume includes the predicted total volume of the inbound and outbound goods and the predicted total quantity of the inbound and outbound goods.

[0034] Step S105: Calculate the coefficient of determination between the base prediction value of the inbound and outbound volume and the actual value in the test data.

[0035] Step S106: Select the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model.

[0036] Step S107: Predict the inbound and outbound volume of the warehouse based on the optimal prediction model.

[0037] In some embodiments of the present application, the inbound and outbound volume of the warehouse includes the total volume of the inbound and outbound goods.

[0038] In some embodiments of the present application, the inbound and outbound volume of the warehouse includes the total quantity of the inbound and outbound goods.

[0039] The present application provides a method for predicting the inbound and outbound volume of a warehouse based on an integrated machine learning with a Bagging structure. By obtaining multiple subsets, training an independent base learner on each subset to obtain multiple base prediction models, and comparing the results of the base prediction models to obtain the optimal prediction model and further obtain the final output, it can effectively reduce the high variance problem that may exist in a single model, that is, overfitting the noise data in a specific training set, and is particularly suitable for the problem of high uncertainty and large volatility of the inbound and outbound volume of the warehouse. Training based on multiple subsets enables the prediction process to have the ability to resist outliers and noise. Even if some subsets contain inaccurate data points, the impact they cause will be offset by the good performance of other subsets.

[0040] The following is a further detailed description, taking "day" as an example of the "first set unit period". In some embodiments of the present application, the original sample data includes: "the total volume of goods shipped out daily", "the total volume of goods received daily", "the total number of goods shipped out daily", "the total number of goods received daily", "the shipping time of goods", and "the receiving time of goods".

[0041] Multiple subsets based on the original sample data include: a first subset, a second subset, a third subset, and a fourth subset.

[0042] The first subset includes: the total volume of goods shipped out of the warehouse daily and the total volume of goods entering the warehouse daily. The first subset is the basic subset, containing the basic information of inbound and outbound.

[0043] The second subset includes: the total volume of goods shipped out of the warehouse daily, the total volume of goods entering the warehouse daily, the total number of goods shipped out of the warehouse daily, and the total number of goods entering the warehouse daily. The second subset adds the quantity information of goods inbound and outbound to enable the basic learner to capture the relationship between volume and quantity.

[0044] The third subset is the data after smoothing the first subset. The smoothing process is used to reduce the noise in the basic subset and help the basic learner capture a more stable trend.

[0045] The fourth subset is generated by the following method: Based on the shipping time or receiving time of goods, and according to a preset time period, group the total volume of goods shipped out of the warehouse daily and the total volume of goods entering the warehouse daily. The fourth subset is used to capture the inbound and outbound patterns in different time periods, enabling the basic learner to understand the impact of time factors on the inbound and outbound quantities.

[0046] When taking "day" as the "first set unit period", the preset time period is in "hours" as the unit. For example, every six hours is a unit, such as "00:00 - 06:00", "06:00 - 12:00", "12:00 - 18:00", "18:00 - 24:00".

[0047] When taking "week" as the "first set unit period", the preset time period is in "days" as the unit, so that the basic learner can understand the impact of weekdays and holidays on the inbound and outbound quantities.

[0048] When taking "year" as the "first set unit period", the preset time period is in "quarters" as the unit, so that the basic learner can understand the impact of seasons on the inbound and outbound quantities.

[0049] In some embodiments of the present application, the first subset may also include: the total quantity of goods shipped out of the warehouse daily and the total quantity of goods entering the warehouse daily. Similarly, the first subset is used as the basic subset. Similarly, the fourth subset groups the total quantity of goods shipped out of the warehouse daily and the total quantity of goods entering the warehouse daily based on the shipping time or receiving time of the goods according to a preset time period.

[0050] The division of the first subset, the second subset, the third subset, and the fourth subset takes into account the diversity and independence of the input features, and at the same time introduces the time factor, enabling the base learner to capture different aspects of the input features and improving the prediction performance of the ensemble machine learning.

[0051] In some embodiments of the present application, the base learner includes a first base learner, a second base learner, a third base learner, and a fourth base learner, where:

[0052] The first base learner is a long short-term neural network, the second base learner is a multi-linear regression algorithm, the third base learner is a support vector machine, and the fourth base learner is a random forest algorithm.

[0053] Based on the above subset generation method, the long short-term neural network can capture long-term dependencies, the multi-linear regression algorithm can capture the linear relationships between input features, especially the linear relationship between volume and quantity. The support vector machine can handle high-dimensional data, and the random forest algorithm can handle complex non-linear relationships. The design of the four base learners can provide diverse prediction perspectives, and can complement each other, adapt to the data characteristics of the subsets, and make full use of their respective advantages.

[0054] In some embodiments of the present application, a dataset partitioning tool is used to generate multiple sets of different training data and test data based on multiple subsets of the original sample data.

[0055] Exemplarily, 90% of the sample data is extracted from the first subset, the second subset, the third subset, and the fourth subset as training data using the train_test_split function, and the remaining 10% of the sample data is used as test data.

[0056] The train_test_split function can be expressed as:

[0057] x_train, x_test, y_train, y_test = train_test_split(train_data, train_target, test_size, random_state, shuffle)

[0058] Among them, x_train, x_test, y_train, and y_test are return values;

[0059] x_train is the feature variable of the training set;

[0060] x_test is the feature variable of the test set;

[0061] y_train is the target variable of the training set;

[0062] y_test is the target variable of the test set.

[0063] train_data is the input feature variable, that is, the independent variable, and it is the sample feature parameter data set to be divided. In this embodiment, it is the input feature variable.

[0064] train_target is the target variable, that is, the dependent variable, and it is the sample target parameter data set to be divided. In this embodiment, it is the total volume of daily inbound and outbound goods.

[0065] test_size is the data volume of the test set; test_size can be a floating-point number between 0 and 1, used to specify the proportion of the test set; test_size can also be an integer, directly representing the number of data in the test set, and the remaining data is included in the training set. For example, in this embodiment, it can be set to 0.1.

[0066] random_state is the random seed, used to ensure the consistency of the results of each split. The data set is randomly shuffled and divided according to the random seed. If different random number seeds are set each time, the train_test_split function will divide the data set into the training set and the test set in different ways. If the same random number seed is set each time, the same training set and test set will be obtained each time. For example, in this embodiment, it can be set to 42.

[0067] shuffle is used to determine whether to randomly reorder the data before dividing the data set. If it is False, the original order of the sample data is not disrupted. If it is True, the order of the sample data is disrupted.

[0068] The train_test_split function method can automate the process of data set division, greatly saving time, and can achieve random division by setting the random seed, thereby avoiding human preferences and ensuring the randomness of the data set.

[0069] Input the training data corresponding to the first subset, the second subset, the third subset, and the fourth subset into four basic learners respectively to form 16 different combinations. When determining the optimal hyperparameter combinations of each basic prediction model, grid search for hyperparameter optimization can be further adopted, and combined with ten-fold cross-validation to evaluate and determine the model performance corresponding to each hyperparameter combination. Use the optimal hyperparameter combinations of each determined basic prediction model to train and test each basic prediction model respectively, and finally output multiple basic predicted values of the inbound and outbound quantities.

[0070] Specifically, in the process of training a pre-configured and independent basic learner with each group of training data to obtain multiple basic prediction models, the following steps can be executed:

[0071] Preset preliminary hyperparameter combinations (for example, different penalty parameters and kernel functions can be selected for the support vector machine according to experience or literature). Further define a hyperparameter grid and use grid search to traverse all possible hyperparameter combinations. For each group of hyperparameter combinations, use ten-fold cross-validation. When using ten-fold cross-validation, divide the training data into 10 subsets. In each round, use 9 subsets for training and 1 subset for validation, repeat 100 times, calculate the average performance of each hyperparameter combination; finally, compare the cross-validation results of all hyperparameter combinations, select the hyperparameter combination with the best performance, use the training data and the selected best hyperparameter combination to train the basic learner, and adjust the actual parameters of the basic learner to obtain the basic prediction model.

[0072] Use each basic prediction model to perform basic predictions on each group of test data, output multiple basic predicted values of the inbound and outbound quantities, and at the same time calculate the performance metrics of the basic prediction models to understand the performance of the basic learners on unseen data.

[0073] Coefficient of determination R 2 The defining formula is:

[0074]

[0075] where RSS is the residual sum of squares, which is the sum of the squares of the differences between the actual values and the predicted values; TSS is the total sum of squares, which is the sum of the squares of the differences between the actual values and their mean values;

[0076]

[0077] where y i is the i-th actual value, is the predicted value corresponding to y i , is the mean value corresponding to y i , and n is the number of sample data. When R 2When R = 1, it indicates that the basic prediction model perfectly fits the original sample data, and the actual values are all on the regression line without residuals; when R 2 = 0, it means that the basic prediction model fails to explain any data variability; when R 2 < 0, it usually means that the basic prediction model is overfitting or there are other serious problems. Therefore, select the basic prediction model with the coefficient of determination R 2 closest to 1 as the optimal prediction model.

[0078] In some embodiments of the present application, preferably, the original sample data includes: within a third set unit period, the total volume of goods warehoused and outwarehoused in each first set unit period, the total quantity of goods warehoused and outwarehoused in each first set unit period, the outwarehousing time of the goods, or the warehousing time of the goods. Predict the warehouse warehousing and outwarehousing volume of the second set unit period based on the optimal prediction model.

[0079] Corresponding to the first set unit period in days, the third set unit period can be 20 days, and the second set unit period can be 3 days. The length of the third set unit period is greater than the length of the second set unit period.

[0080] In some embodiments of the present application, the method for predicting the warehouse warehousing and outwarehousing volume based on integrated machine learning further includes a calibration step, specifically as Figure 2 shown.

[0081] Step S201: Predict the warehouse warehousing and outwarehousing volume of the second set unit period based on the optimal prediction model.

[0082] Step S202: Set a time point within the second set unit period, and obtain the period sample data from the starting point of the second set unit period to the set time point from the warehouse side.

[0083] Step S203: Add the period sample data to the original sample data to generate calibrated sample data.

[0084] The second set unit period is 3 days. At 19:00 on the first day, the period sample data from the starting point of the second set unit period to 19:00 on the first day, that is, the real data of the first day, can be added to the original sample data to generate calibrated sample data.

[0085] Step S204: Obtain multiple subsets of the calibrated sample data, and divide each subset into calibrated training data and calibrated test data.

[0086] Step S205: Use each group of calibrated training data to train multiple basic learners to update multiple basic prediction models.

[0087] Step S206: Use each updated basic prediction model to perform basic prediction on the corrected test data, and output multiple updated inbound and outbound volume basic prediction values.

[0088] Step S207: Calculate the coefficient of determination between the basic prediction value of the inbound and outbound volume and the actual value in the test data.

[0089] Step S208: Select again the basic prediction model with the coefficient of determination closest to 1 as the optimal prediction model.

[0090] Step S209: Based on the optimally predicted model selected again, predict the inbound and outbound volume of the warehouse in the second set unit period again.

[0091] By setting time points within the second set unit period, obtaining the latest period sample data, and adding it to the original sample data to generate corrected sample data, this method can adapt to data changes, ensure that the optimally predicted model is always selected, improve the generalization ability of the model, reduce the risk of overfitting, and the optimally predicted model will have better performance when facing new data. The second set unit period and the third set unit period can be set according to business requirements within the first set unit period, flexibly adjusting the prediction range and accuracy of the optimally predicted model. The long-term data of the third set unit period is used for training, and the second set unit period is used for prediction, and the short-term data among them is used for correction. This combination can find a balance between long-term trends and short-term fluctuations, that is, reduce the sensitivity of the model to short-term noise and accurately reflect the current trend changes.

[0092] Since training is performed based on multiple different subsets and basic learners, the multi-core processor or distributed computing resources can be fully utilized to execute the warehouse inbound and outbound volume prediction method provided by this application based on integrated machine learning, thereby improving the modeling speed. In the face of large-scale historical data, it has better efficiency. Compared with the prior art, this application does not require complex tuning of hyperparameters, has simple data requirements, does not rely on additional feature engineering or external variables, and can be used only after determining the subset and the basic learner, greatly reducing the usage threshold and can also be popularized and used in small and medium-sized warehouses without professional personnel.

[0093] As Figure 3 and Figure 4 shown, the second aspect of this application provides a warehouse inbound and outbound volume prediction system 1 based on integrated machine learning, including:

[0094] A processing device 10, which includes:

[0095] The first acquisition module 101 is configured to acquire original sample data from the warehouse side; the original sample data at least includes: the total volume of goods warehoused and outwarehoused in the first set unit cycle, the total quantity of goods warehoused and outwarehoused in the first set unit cycle, the outbound time of the goods or the inbound time of the goods;

[0096] The partitioning module 102 is configured to acquire multiple subsets of the original sample data and partition each subset into training data and test data;

[0097] The training module 103 is configured to train a pre-configured and independent base learner using each set of training data to obtain multiple base prediction models;

[0098] The base prediction module 104 is configured to perform base prediction on each set of test data using each base prediction model and output multiple base predicted values of the warehousing and outwarehousing quantities; wherein, the base predicted values of the warehousing and outwarehousing quantities include the predicted total volume of the warehoused and outwarehoused goods and / or the predicted total quantity of the warehoused and outwarehoused goods;

[0099] The verification module 105 is configured to calculate the coefficient of determination between the base predicted values of the warehousing and outwarehousing quantities and the actual values in the test data;

[0100] The selection module 106 is configured to select the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model;

[0101] The actual prediction module 107 is configured to predict the warehouse warehousing and outwarehousing quantities based on the optimal prediction model; the warehouse warehousing and outwarehousing quantities include the total volume of the warehoused and outwarehoused goods and / or the total quantity of the warehoused and outwarehoused goods.

[0102] In this application, the processing device 10 may be a server or a combination of one or more high-performance computers. The processing device 10 is composed of a processor, a storage unit (including volatile memory and non-volatile memory), a display device, an operating device, a communication interface, a driving device, etc., and is connected through a bus. The processor may be a dedicated processor, a central processing unit (CPU), etc. The processor can access the storage unit and execute instructions or application programs stored therein to complete related functions. The display device is used to display various information. The operating device is used to receive user operation instructions, interact with the storage unit or storage medium, and process interrupt signals. In one or more embodiments of this application, the storage medium may be a CD-ROM, a floppy disk, an optical magneto-optical disc, etc., a medium that records data by optical, electrical or magnetic means. The storage medium may also be a semiconductor memory such as a ROM, a flash memory, etc. that stores information by electrical means.

[0103] Each of the modules included in the processing device 10 can be implemented by the processor running a program.

[0104] In some embodiments of the present application, the original sample data includes: within a third set unit period, the total volume of goods warehoused and unwarehoused in each first set unit period, the total quantity of goods warehoused and unwarehoused in each first set unit period, the outbound time of the goods or the inbound time of the goods;

[0105] The actual prediction module 107 is configured to predict the inbound and outbound volume of the warehouse in a second set unit period based on the optimal prediction model;

[0106] The prediction system 1 further includes:

[0107] The second acquisition module 201 is configured to obtain the period sample data from the starting point of the second set unit period to the set time point from the warehouse end at a set time point within the second set unit period;

[0108] The update module 202 is configured to add the period sample data to the original sample data to generate corrected sample data;

[0109] The correction division module 203 is configured to obtain multiple subsets of the corrected sample data and divide each subset into corrected training data and corrected test data;

[0110] The correction training module 204 is configured to train multiple base learners using each group of corrected training data to update multiple base prediction models;

[0111] The corrected base prediction module 205 is configured to perform base prediction on the corrected test data using each updated base prediction model and output multiple updated inbound and outbound volume base prediction values;

[0112] The correction verification module 206 is configured to calculate the coefficient of determination between the base prediction value of the inbound and outbound volume and the actual value in the test data;

[0113] The correction selection module 207 is configured to reselect the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model;

[0114] The corrected actual prediction module 208 is configured to predict the inbound and outbound volume of the warehouse in the second set unit period again based on the reselected optimal prediction model;

[0115] Wherein, the first set unit period is longer than the second set unit period.

[0116] In some embodiments of the present application, in the division module 102, the multiple subsets of the original sample data include:

[0117] The first subset, which includes the volume of goods outbound from the warehouse within the first set unit period and the volume of goods entering the warehouse within the first set unit period;

[0118] The second subset includes the volume of goods shipped out of the warehouse within the first set unit period, the volume of goods entering the warehouse within the first set unit period, the quantity of goods shipped out of the warehouse within the first set unit period, and the quantity of goods entering the warehouse within the first set unit period;

[0119] The third subset is generated after smoothing the first subset; and

[0120] The fourth subset is generated by the following method: based on the outbound time or inbound time of the goods, and according to a preset time period, the volume of goods shipped out of the warehouse within the first set unit period and the total volume of goods entering the warehouse within the first set unit period are grouped.

[0121] In some embodiments of the present application, the first basic learner is a long short-term neural network, the second basic learner is a multi-linear regression algorithm, the third basic learner is a support vector machine, and the fourth basic learner is a random forest algorithm.

[0122] In some embodiments of the present application, a dataset partitioning tool is used to generate multiple sets of different training data and test data based on multiple subsets of the original sample data.

[0123] The prediction system 1 provided by the present application has the same technical effects and will not be elaborated here.

[0124] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions required to be protected by the present invention.

Claims

1. A method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning, characterized in that, Including: Obtaining original sample data from the warehouse side; The original sample data at least includes: the total volume of goods warehoused and unwarehoused in the first set unit cycle, the total quantity of goods warehoused and unwarehoused in the first set unit cycle, the outbound time of the goods or the inbound time of the goods; Obtaining multiple subsets of the original sample data, and dividing each subset into training data and test data; Training a pre-configured and independent base learner with each group of training data to obtain multiple base prediction models; Using each of the base prediction models to perform base prediction on each group of the test data, and outputting multiple basic predicted values of the warehousing and outbound quantity; wherein, the basic predicted value of the warehousing and outbound quantity includes the predicted total volume of the warehoused and outbound goods and / or the predicted total quantity of the warehoused and outbound goods; Calculating the coefficient of determination between the basic predicted value of the warehousing and outbound quantity and the actual value in the test data; Selecting the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model; Predicting the warehouse warehousing and outbound quantity based on the optimal prediction model; the warehouse warehousing and outbound quantity includes the total volume of the warehoused and outbound goods and / or the total quantity of the warehoused and outbound goods.

2. The method for predicting the warehouse warehousing and outbound quantity based on integrated machine learning according to claim 1, wherein The original sample data includes: within the third set unit cycle, the total volume of goods warehoused and unwarehoused in each first set unit cycle, the total quantity of goods warehoused and unwarehoused in each first set unit cycle, the outbound time of the goods or the inbound time of the goods; Predicting the warehouse warehousing and outbound quantity in the second set unit cycle based on the optimal prediction model; Setting a time point within the second set unit cycle, and obtaining the period sample data from the starting point to the set time point of the second set unit cycle from the warehouse side; Adding the period sample data to the original sample data to generate corrected sample data; Obtaining multiple subsets of the corrected sample data, and dividing each subset into corrected training data and corrected test data; Training multiple of the base learners with each group of corrected training data to update multiple base prediction models; Using each updated base prediction model to perform base prediction on the corrected test data, and outputting multiple updated basic predicted values of the warehousing and outbound quantity; Calculating the coefficient of determination between the basic predicted value of the warehousing and outbound quantity and the actual value in the test data; Selecting again the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model; Predicting again the warehouse warehousing and outbound quantity in the second set unit cycle based on the optimally selected prediction model again; Wherein, the third set unit cycle is longer than the second set unit cycle.

3. The method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning according to claim 2, wherein The multiple subsets based on the original sample data include: The first subset, the first subset includes the volume of goods outbound from the warehouse within the first set unit cycle, and the volume of goods entering the warehouse within the first set unit cycle; The second subset, the second subset includes the volume of goods outbound from the warehouse within the first set unit cycle, the volume of goods entering the warehouse within the first set unit cycle, the quantity of goods outbound from the warehouse within the first set unit cycle, and the quantity of goods entering the warehouse within the first set unit cycle; A third subset, generated after smoothing the first subset; and A fourth subset, generated by the following method: based on the outbound time or inbound time of the goods, grouping the volume of goods outbound from the warehouse within the first set unit cycle and the total volume of goods entering the warehouse within the first set unit cycle according to a preset time period.

4. The method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning according to claim 3, wherein The basic learners include a first basic learner, a second basic learner, a third basic learner, and a fourth basic learner, where: The first basic learner is a long short-term neural network, the second basic learner is a multi-linear regression algorithm, the third basic learner is a support vector machine, and the fourth basic learner is a random forest algorithm.

5. The method for predicting the inbound and outbound volume of a warehouse based on integrated machine learning according to claim 4, wherein Using a dataset partitioning tool, multiple sets of different training data and test data are generated based on multiple subsets of the original sample data.

6. A warehouse inbound and outbound volume prediction system based on integrated machine learning, characterized in that Including: A processing device, which includes: A first acquisition module configured to acquire original sample data from the warehouse side; the original sample data at least includes: the total volume of goods inbound and outbound within the first set unit cycle, the total quantity of goods inbound and outbound within the first set unit cycle, the outbound time or inbound time of the goods. A partitioning module configured to acquire multiple subsets of the original sample data and partition each subset into training data and test data. A training module configured to train a pre-configured and independent basic learner using each set of training data to obtain multiple basic prediction models. A basic prediction module configured to perform basic prediction on each set of the test data using each of the basic prediction models and output multiple basic predicted values of the inbound and outbound volume; wherein, the basic predicted values of the inbound and outbound volume include the predicted total volume of inbound and outbound goods and / or the predicted total quantity of inbound and outbound goods. A verification module configured to calculate the coefficient of determination between the basic predicted values of the inbound and outbound volume and the actual values in the test data. A selection module configured to select the basic prediction model with the coefficient of determination closest to 1 as the optimal prediction model. An actual prediction module configured to predict the inbound and outbound volume of the warehouse based on the optimal prediction model; the inbound and outbound volume of the warehouse includes the total volume of inbound and outbound goods and / or the total quantity of inbound and outbound goods.

7. The system for predicting the inbound and outbound volume of a warehouse based on integrated machine learning according to claim 6, wherein The original sample data includes: within the third set unit cycle, the total volume of goods inbound and outbound in each first set unit cycle, the total quantity of goods inbound and outbound in each first set unit cycle, the outbound time or inbound time of the goods. The actual prediction module is configured to predict the inbound and outbound volume of the warehouse in the second set unit cycle based on the optimal prediction model. The prediction system further includes: A second acquisition module configured to acquire the period sample data from the starting point of the second set unit cycle to the set time point from the warehouse side at a set time point within the second set unit cycle. An update module configured to add the period sample data to the original sample data to generate corrected sample data; A correction division module configured to obtain multiple subsets of the corrected sample data and divide each subset into corrected training data and corrected test data; A correction training module configured to train multiple of the base learners using each set of corrected training data to update multiple base prediction models; A corrected base prediction module configured to perform base prediction on the corrected test data using each updated base prediction model and output multiple updated inbound / outbound volume base prediction values; A correction verification module configured to calculate the coefficient of determination between the inbound / outbound volume base prediction values and the actual values in the test data; A correction selection module configured to re-select the base prediction model with the coefficient of determination closest to 1 as the optimal prediction model; A corrected actual prediction module configured to re-predict the warehouse inbound / outbound volume for a second set unit period based on the re-selected optimal prediction model; wherein the first set unit period is longer than the second set unit period.

8. The warehouse inbound and outbound volume prediction system based on integrated machine learning according to claim 7, wherein In the division module, the multiple subsets of the original sample data include: A first subset including the volume of goods shipped out of the warehouse within a first set unit period and the volume of goods entering the warehouse within the first set unit period; A second subset including the volume of goods shipped out of the warehouse within a first set unit period, the volume of goods entering the warehouse within the first set unit period, the quantity of goods shipped out of the warehouse within the first set unit period, and the quantity of goods entering the warehouse within the first set unit period; A third subset generated by smoothing the first subset; and A fourth subset generated by the following method: grouping the volume of goods shipped out of the warehouse within a first set unit period and the total volume of goods entering the warehouse within the first set unit period according to a preset time period based on the outbound time or inbound time of the goods.

9. The warehouse inbound / outbound volume prediction system based on ensemble machine learning according to claim 8, wherein The base learners include a first base learner, a second base learner, a third base learner, and a fourth base learner, wherein: The first base learner is a long short-term neural network, the second base learner is a multi-linear regression algorithm, the third base learner is a support vector machine, and the fourth base learner is a random forest algorithm.

10. The warehouse inbound / outbound volume prediction system based on ensemble machine learning according to claim 9, wherein Using a dataset division tool, multiple sets of different training data and test data are generated based on the multiple subsets of the original sample data.