Parcel quantity prediction method and device, equipment, storage medium and program product

By acquiring and preprocessing parcel data in logistics business, combining parcel volume prediction model, considering historical and planned data, the accuracy of parcel volume prediction is solved and efficient logistics business arrangements are achieved.

CN120494882APending Publication Date: 2025-08-15BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510559592.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

It is difficult to accurately predict the volume of parcels in logistics business in the prior art, especially during business growth stages and special time nodes, which leads to difficulties in forecasting.

Method used

By obtaining the first time sequence data within the first historical time and the parcel planning data of the first prediction time, pre-processing is performed to make predictions using the parcel quantity prediction model, considering the planned data of the historical actual parcel quantity and prediction time, and quickly adjusting the prediction results.

Benefits of technology

It improves the accuracy of parcel volume forecasting and meets the efficient production capacity arrangements of logistics business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494882A_ABST
    Figure CN120494882A_ABST
Patent Text Reader

Abstract

The invention discloses a parcel quantity prediction method and device, equipment, a storage medium and a program product, and relates to the fields of logistics technology, artificial intelligence technology, large model technology and large language model.The method comprises the steps that first time sequence data of a first parcel in a first historical duration and parcel plan data of first prediction time are acquired, the first time sequence data is used for indicating a relationship between the time node and package information, the package information comprises an actual package amount and region information, and the region information is used for indicating a region to which a resource acquisition space corresponding to the package belongs; preprocessing the first time sequence data to obtain second time sequence data; and utilizing a package quantity prediction model to predict the package quantity at the first prediction time to obtain a package quantity prediction result of the region at the first prediction time, the package quantity prediction model being configured to be used for outputting the package quantity prediction result based on the second time sequence data and the package plan data at the first prediction time. The method can improve the accuracy of package quantity prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of logistics technology, artificial intelligence technology, large model technology, and large language model technology, and specifically to methods, devices, equipment, storage media, and program products for predicting package volume. Background Art

[0002] In the logistics business, accurately forecasting daily package volume is crucial for warehouse staffing. This is especially true during growth phases, when package volume fluctuates significantly. For example, at specific time points, package volume can experience significant sudden changes. Consequently, these dramatic fluctuations in actual business data can significantly complicate accurate package volume forecasting. Summary of the Invention

[0003] In view of this, the present disclosure provides a package volume prediction method, apparatus, device, storage medium, and program product to solve the problem of accurate package volume prediction.

[0004] In a first aspect, the present disclosure provides a method for predicting package volume, the method comprising:

[0005] Obtaining first time series data of a first package within a first historical time period and package plan data for a first predicted time period, where the first time series data indicates a relationship between a time node and package information, and the package information includes actual package quantity and region information, where the region information indicates a region to which a resource acquisition space corresponding to the package belongs;

[0006] Preprocessing the first time series data to obtain second time series data;

[0007] The package volume prediction model is used to predict the package volume at the first prediction time to obtain the package volume prediction result of the area at the first prediction time. The package volume prediction model is configured to output the package volume prediction result based on the second time series data and the package plan data at the first prediction time.

[0008] In a second aspect, the present disclosure provides a package volume prediction device, the device comprising:

[0009] an acquisition module, configured to acquire first time series data of a first package within a first historical time period and package plan data for a first predicted time period, wherein the first time series data indicates a relationship between time and package information, the package information includes an actual package quantity and region information, and the region information indicates a region to which a resource acquisition space corresponding to the package belongs;

[0010] a preprocessing module, configured to preprocess the first time series data to obtain second time series data;

[0011] A prediction module is used to use a package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time, and the package volume prediction model is configured to output the package volume prediction result based on the second time series data and the package plan data at the first prediction time.

[0012] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the package quantity prediction method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the package quantity prediction method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0014] In a fifth aspect, the present disclosure provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the package quantity prediction method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0015] The package volume prediction method provided by the disclosed embodiments obtains first time series data within a first historical time period and package plan data for a first prediction time period, preprocesses the first time series data to obtain second time series data, and uses a package volume prediction model to predict the package volume at the first prediction time period to obtain a package volume prediction result for a region at the first prediction time period. The first time series data indicates the relationship between time nodes and package information, and the package information includes actual package volume and regional information. Specifically, the actual package volume and regional information for each time node within the first historical time period are obtained. The preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization, ensuring the diversity and discretization of the second time series data, providing favorable input data for subsequent processing of the package volume prediction model. When using the package volume prediction model to predict package volume, not only the historical actual package volume but also the planned data for the prediction time period are considered. Furthermore, the prediction results can be quickly adjusted based on recent data changes, improving the accuracy of the package volume prediction, thereby meeting the efficient capacity scheduling requirements of logistics services. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;

[0018] Figure 2 is a flowchart of a package volume prediction method according to an embodiment of the present disclosure;

[0019] Figure 3 is a flowchart of another package volume prediction method according to an embodiment of the present disclosure;

[0020] Figure 4 is a flowchart of another package volume prediction method according to an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of the training of the package volume prediction model and the package volume prediction according to an embodiment of the present disclosure;

[0022] Figure 6 is a structural block diagram of a package volume prediction device according to an embodiment of the present disclosure;

[0023] Figure 7 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0030] In related technologies, time series forecasting technology is usually used to predict the volume of logistics packages, including but not limited to moving average (MA), autoregressive integrated sliding average model plus exogenous variables (arimax), Prophet model (a time series forecasting model designed for processing time series data with strong seasonality, trend changes, missing values and outliers), Autoformer model (a deep learning model for processing time series forecasting tasks), Informer model (a Transformer model designed for long-series time series forecasting), PatchTST model (a time series forecasting method based on Transformer), etc.

[0031] Among them, the advantage of the moving average (MA) method is its simplicity, but its disadvantage is that it only predicts a single variable and cannot accurately predict the peak package volume during special periods such as big sales.

[0032] The advantage of the autoregressive integrated moving average model plus exogenous variables (ARIMAX) model is its ability to predict multiple variables. Its disadvantage is that it has strict requirements on data, requiring the data to be stationary. However, business data in the growth period and with frequent promotional activities usually does not meet the stationary conditions, which limits the application of ARIMAX.

[0033] The advantages of the Prophet model are that it can capture trends, seasons, and other information in time series data and use future known covariates to assist in prediction. However, its disadvantages are that it is more suitable for long-term predictions and is insensitive to short-term data trend changes. When the logistics package volume fluctuates rapidly in the short term, the prediction effect is poor.

[0034] The advantage of the Autoformer model is that it can adaptively learn the trends and periodicity of time series data. However, its disadvantage is that it cannot use future known covariates and multivariate predictions, making it difficult to meet the forecasting needs of current logistics business.

[0035] The advantages of the Informer model are that it can significantly reduce computing resources and training time through the sparse self-attention mechanism and has the ability to handle multiple variables. However, its disadvantages are that it cannot use future known variables and cannot distinguish between primary variables and covariates. In practical applications, it is difficult to distinguish the impact of each variable on package volume prediction.

[0036] The advantage of the PatchTST model is that it can reduce computational complexity by dividing the time series into patches while retaining local information in the sequence, and better learn the characteristic patterns in the time series. However, its disadvantage is that it cannot use future known variables and cannot learn the correlation between multiple variables, making it difficult to handle the mutual influence of multiple factors in the logistics package volume forecast.

[0037] Based on this, the package volume prediction method provided by the embodiment of the present disclosure performs package volume prediction based on package planning data, prediction time and first time series data within a first historical period. It not only takes into account the actual historical package volume, but also takes into account the planned data at the prediction time. At the same time, it can quickly adjust the prediction results according to recent data changes, thereby improving the accuracy of package volume prediction, thereby meeting the efficient production capacity arrangement of logistics business.

[0038] As an optional application scenario of the embodiment of the present disclosure, Figure 1 As shown, the system includes a terminal 101 and a server 102. Server 102 executes the package volume prediction method provided in the embodiments of the present disclosure to obtain a package volume prediction result at a first prediction time, and feeds the obtained package volume prediction result back to terminal 101. The logistics dispatcher can obtain the package volume prediction result at the first prediction time through terminal 101 and perform logistics dispatch based on the package volume prediction result.

[0039] Of course, the package volume prediction results can also have other downstream applications, which are not limited here and will be determined based on the actual application scenario.

[0040] According to an embodiment of the present disclosure, an embodiment of a package volume prediction method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0041] In this embodiment, a package volume prediction method is provided, which can be used in logistics scenarios. Figure 2 is a flowchart of a method for predicting package volume according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps:

[0042] Step S201: Acquire first time series data within a first historical period and package planning data at a first predicted time.

[0043] Among them, the first time series data is used to indicate the relationship between the time node and the package information. The package information includes the actual package quantity and regional information. The regional information is used to indicate the region to which the resource acquisition space corresponding to the package belongs.

[0044] The first historical duration is used to indicate the period from the current time. This period can be 14 days, 10 days, or 7 days, etc., and is set according to actual needs. There is no limitation here. For example, if the current time is April 15, the first historical duration is used to indicate the period from April 1 to April 14.

[0045] The first time series data is used to indicate the relationship between a time node and package information. The time node can be every day, and the package information includes the actual package volume and regional information. The actual package volume can be used to represent the actual number of packages generated each day, and the regional information is used to indicate the region to which the resource acquisition space corresponding to the package belongs. Specifically, the resource acquisition space is used to provide item resources. Item resources need to be transported from location A to location B in the form of packages. Since there are many item resources in the same resource acquisition space, and not all item resources are necessarily transported from location A, in order to simplify the complexity of data statistics, the package volume is counted from the dimension of the resource acquisition space.

[0046] The same region can contain multiple resource acquisition spaces, and the same resource acquisition space can provide multiple item resources. The package information in the first time series can represent the region to which the corresponding resource acquisition space belongs. For example, if item resources 1 through 10 all come from resource acquisition space a1, which belongs to region c1, then the region information for the packages corresponding to item resources 1 through 10 is region c1, to which resource acquisition space a1 belongs.

[0047] The first time series data within the first historical period indicates the relationship between the time node and the package information. Taking time node 1 as an example, the representation of the time series data can be:

[0048] Time node 1 - package p1 - resource acquisition space a1 - area c1;

[0049] Time node 1 - package p2 - resource acquisition space a1 - area c1;

[0050] Time node 1 - package p3 - resource acquisition space a2 - area c1;

[0051] Time node 1 - package p4 - resource acquisition space a3 - area c2.

[0052] From the above example, it can be seen that at time node 1, there are corresponding p1 to p3, which come from resource acquisition spaces a1 to resource acquisition spaces a3, and correspondingly, areas c1 and c2.

[0053] Of course, the data dimensions included in the first time series data are not limited to those shown above, and may also include other data dimensions, which are not limited here.

[0054] The first prediction time is used to indicate the predicted time node. For example, if the current time is April 15, the first prediction time includes April 16 to April 26. Of course, the prediction period corresponding to the first prediction time is not limited here and can be 14 days, 10 days, etc.

[0055] The package plan data for the first predicted time period can be based on time nodes, with package plan data for each time node being given. The package plan data for the first predicted time period can be obtained from the resource recommendation platform. Before the resource acquisition space recommends resources on the resource recommendation platform, the package plan data can be configured on the resource recommendation platform. Accordingly, the package plan data can be obtained. Of course, other methods can also be used to obtain the package plan data, and this is not limited to any method herein.

[0056] Step S202: preprocess the first time series data to obtain second time series data.

[0057] The preprocessing includes at least one of time series data division, time series data organization and time series data discretization.

[0058] Since the first time series data includes many data dimensions, if only the regional dimension is used to divide the time series data, the amount of time series data may be small. Therefore, the dimension can be further refined based on the regional dimension, that is, the time series data can be divided with a finer granularity. For example, it can be refined to the dimension of the resource acquisition space, or the dimension of the resource provider to which the resource acquisition space belongs, etc. Among them, the same resource provider may have multiple resource acquisition spaces, and these multiple resource acquisition spaces can be in the same region or in different regions.

[0059] Time series data organization includes data filling and data validity screening in time series data. For example, if there is no package volume in the resource acquisition space at a certain time node, the package volume at that time node can be filled with 0 to ensure time continuity and consistency of data length.

[0060] Time series data discretization is used to discretize time series data to ensure that the processed time series data can be used to perform time series regression tasks using a large language model.

[0061] Of course, the preprocessing of the first time series data is not limited to what is shown above, and may also include other processing methods, which are not limited here.

[0062] Step S203: Use the package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time.

[0063] The package volume prediction model is configured to output a package volume prediction result based on the second time series data and the package planning data of the first prediction time.

[0064] The package volume prediction model's input is configured to include the second time series data and the package plan data for the first prediction time, and its output is configured to be the package volume prediction result. The package volume prediction model is constructed based on the large prediction model, and its specific base model is not limited. For example, a base model that is not manually aligned can be selected.

[0065] The package volume prediction method provided in this embodiment obtains first time series data within a first historical time period and package plan data for a first prediction time period, preprocesses the first time series data to obtain second time series data, and uses a package volume prediction model to predict the package volume at the first prediction time period to obtain a package volume prediction result for a region at the first prediction time period. The first time series data indicates the relationship between time nodes and package information, and the package information includes actual package volume and regional information. Specifically, the actual package volume and regional information for each time node within the first historical time period are obtained. The preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization, ensuring the diversity and discretization of the second time series data, providing favorable input data for subsequent processing of the package volume prediction model. When using the package volume prediction model to predict package volume, not only the historical actual package volume but also the planned data for the prediction time period are considered. Furthermore, the prediction results can be quickly adjusted based on recent data changes, improving the accuracy of the package volume prediction, thereby meeting the requirements for efficient capacity scheduling in logistics operations.

[0066] In this embodiment, a package volume prediction method is provided, which can be used in logistics scenarios. Figure 3 is a flowchart of a method for predicting package volume according to an embodiment of the present disclosure. Figure 3 As shown, the process includes the following steps:

[0067] Step S301: Acquire first time series data within a first historical period and package planning data at a first predicted time.

[0068] The first time series data is used to indicate the relationship between the time node and the package information. The package information includes the actual package quantity and the regional information. The regional information is used to indicate the region to which the resource acquisition space corresponding to the package belongs. Figure 2 The step S201 of the illustrated embodiment is not limited in any way herein.

[0069] Step S302: preprocess the first time series data to obtain second time series data.

[0070] The preprocessing includes at least one of time series data division, time series data organization and time series data discretization.

[0071] Exemplarily, the above step S302 includes:

[0072] Step S3021: Based on the resource acquisition space in the region information, the first time series data is bucketed to obtain data in the time series bucket.

[0073] The data in the time series bucket includes time series sub-data of multiple resource acquisition spaces in the same area.

[0074] As described above, the resource acquisition space has a region to which it belongs. Therefore, the first time series data is dimensionally refined based on the dimension of the resource acquisition space, that is, the time series sub-data of multiple resource acquisition spaces in the same region are divided into the same time series bucket.

[0075] The number of resource acquisition spaces included in the same timing bucket is set according to actual needs. For example, if the same area includes more resource acquisition spaces, the resource acquisition spaces of the same area can be divided into different timing buckets, that is, the same area corresponds to multiple timing buckets; if the same area includes fewer resource acquisition spaces, the resource acquisition spaces of the same area can be divided into the same timing bucket.

[0076] Bucketing of the first time series data may also be related to other factors, which are not listed here one by one.

[0077] The first time series data is a collective term for all time series data within the first historical period. The time series data corresponding to the resource acquisition space within the same region comes from the first time series data. For ease of description, this is referred to as time series sub-data. It should be noted that time series sub-data is a portion of the first time series data, and can also be understood as a data segment within the first time series data.

[0078] In some optional implementations, step S3021 includes:

[0079] Step a1: Determine the resource acquisition space and the number of time series buckets in the same area in the first time series data.

[0080] Step a2: bucketing the first time series data based on the resource acquisition space and the number of time series buckets in the same area to obtain data in the time series buckets.

[0081] The number of time series buckets may be pre-set or obtained by adaptively adjusting the data volume of the first time series data. The specific number is not limited herein.

[0082] For the resource acquisition space within the same region in the first time series data, the time series data in the first time series data can be first divided into regions, and then the resource acquisition space is divided for the time series data in the same region, and finally the time series data of the resource acquisition space in the same region is obtained.

[0083] The first time series data is bucketed based on the number of time series buckets and the resource acquisition space in the same region. That is, the time series data of at least one resource acquisition space in the same region is divided into the same time series bucket to obtain the data in the time series bucket.

[0084] The data in the first time series bucket is bucketed according to the number of time series buckets and the resource acquisition space in the same area, ensuring that the data in the same time series bucket is the time series sub-data of multiple resource acquisition spaces in the same area. The granularity is fine and can ensure the diversity of book sequence data.

[0085] Step S3022: Determine the lifecycle of the time series bucket based on the data in the time series bucket.

[0086] In the same time series bucket, the same resource acquisition space may generate multiple time series data within the first historical duration. Some resource acquisition spaces may be shut down at a certain time point within the first historical duration, while some resource acquisition spaces may be started at a certain time point within the first historical duration. Therefore, for the time series data in the same time series bucket, although the limited duration is the first historical duration, the different shutdown times or start times of the resource acquisition spaces may bring complexity to the time series data in the time series bucket from the time node dimension. Based on this, it is necessary to determine the life cycle of the time series bucket and, on this basis, organize the time series data in the time series bucket.

[0087] The lifecycle of a time series bucket indicates the duration between the start and end of the lifecycle. The lifecycle of a time series bucket is determined based on the data within the bucket. For example, it is determined based on the earliest and latest times when a package volume was generated within the bucket.

[0088] In some optional implementations, the above step S3022 includes:

[0089] Step b1: For the data in the time series bucket, obtain the first time when the package quantity first exists and the second time when the package quantity last exists.

[0090] In step b2, the first time is used as the start time of the life cycle, and the second time is used as the end time of the life cycle.

[0091] For data within a time series bucket, you can query the time when a package quantity first appeared in each resource acquisition space to obtain the time when the package quantity last appeared. Then, based on the resource acquisition spaces within the time series bucket, compare the time when the package quantity first appeared in each resource acquisition space. The earliest time is used as the time when the package quantity first appeared in the time series bucket, recorded as the first time. Compare the time when the package quantity last appeared in each resource acquisition space, and use the latest time as the time when the package quantity last appeared in the time series bucket, recorded as the second time.

[0092] Since the first time and the second time correspond to the time series bucket, the first time is used as the start time of the life cycle, and the second time is used as the end time of the life cycle.

[0093] The life cycle is determined based on the first time when the package quantity first exists and the second time when the package quantity last exists, ensuring that the life cycle is determined based on the actual time when the package occurs, thereby ensuring the accuracy of the life cycle of the time series bucket.

[0094] Step S3023: Process the data in the time series bucket based on the life cycle to obtain second time series sub-data of the time series bucket.

[0095] The second time series data includes the second time series sub-data of all time series buckets.

[0096] After determining the lifecycle of a time series bucket, the time series data for the bucket represents the package information within the continuous duration of the lifecycle. Because some resource acquisition spaces may not generate any packages at a certain point in time, the package volume at that point in time needs to be padded with zeros.

[0097] Furthermore, the time series data within a time series bucket corresponds to multiple resource acquisition spaces. To perform subsequent predictions based on the time series bucket dimension, the time series data from multiple resource acquisition spaces can be fused based on the time node dimension to obtain the second time series sub-data for the time series bucket. The second time series sub-data for all time series buckets are collectively referred to as the second time series data.

[0098] In some optional implementations, step S3023 includes:

[0099] In step c1, the time series sub-data of the resource acquisition space in the time series bucket is supplemented with the package quantity of the entire life cycle to obtain the third time series sub-data of the resource acquisition space.

[0100] In step c2, the third time series sub-data of the space of all resources in the time series bucket is acquired, and the package quantities at the same time node are merged to obtain the second time series sub-data of the time series bucket.

[0101] For the time series sub-data of each resource acquisition space within the same time series bucket, the package volume at consecutive time nodes within the entire life cycle is padded. For example, if the life cycle of a time series bucket is from April 1st to April 14th, and a resource acquisition space has no package volume at the time node of April 8th, the package volume of this resource acquisition space at that time node will be padded to zero.

[0102] To distinguish it from the time series sub-data before padding, the padded time series sub-data is referred to as the third time series sub-data. For all resource acquisition spaces within a time series bucket, the package quantities at the same time node are merged to obtain the total package quantity of the time series bucket at that time node, and accordingly, the second time series sub-data of the time series bucket is obtained.

[0103] The package volume of the entire life cycle of the time series sub-data of the resource acquisition space in the time series bucket is supplemented to facilitate subsequent merging based on time nodes. In addition, the package volume of the third time series sub-data of all resource acquisition spaces in the time series bucket at the same time node is merged to balance the diversity of the time series data and the stability of the target variable.

[0104] In some optional implementations, the above step c2 includes:

[0105] In step c21, for all third time series sub-data in the time series bucket, the package quantity at the same time node is accumulated to obtain fourth time series sub-data.

[0106] Step c22: performing feature discretization processing on the fourth time series sub-data to obtain second time series sub-data.

[0107] For the third time series sub-data of all resource acquisition spaces for the same time series bucket, the package volume at the same time point is accumulated to obtain the total package volume at that time point. Similarly, the fourth time series sub-data of the time series bucket is obtained.

[0108] Because the package volume corresponding to different time nodes may vary significantly, for example, if some specified time nodes are included, the package volume will increase significantly. Therefore, after obtaining the fourth time series sub-data, it is discretized to prevent overfitting in the subsequent model processing. The time series sub-data after the feature discretization processing of the fourth time series sub-data is the second time series sub-data.

[0109] Exemplarily, feature discretization includes but is not limited to equal-frequency discretization, equal-distance discretization, or step discretization, etc. The specific feature discretization method adopted is determined according to actual needs and is not limited here.

[0110] After accumulating the package volume at the same time node, the fourth time series sub-data is obtained. On this basis, the fourth time series sub-data is subjected to feature discretization processing to facilitate the subsequent use of the package volume prediction model to perform time series regression tasks.

[0111] In some optional implementations, the above step c22 includes:

[0112] Step c221: determine the initial discrete coefficient and the minimum value of the package quantity in the fourth time series sub-data, and use the minimum value as the reference value.

[0113] Step c222: Determine a modified dispersion coefficient based on the relationship between the planned package quantity and the reference value corresponding to the time node in the fourth time series sub-data.

[0114] Step c223: Correct the reference value based on the initial discrete coefficient and the corrected discrete coefficient to obtain the equidistant parameters of the time nodes.

[0115] Step c224: Process the package quantity corresponding to the time node based on the equidistant parameter to obtain the second time series sub-data.

[0116] In this embodiment, a stepwise discretization method is used to perform feature discretization processing, wherein the stepwise discretization is related to an equidistant parameter, and the equidistant parameters corresponding to different time nodes may be different.

[0117] After completing and accumulating the data throughout the entire lifecycle, the package volume corresponding to consecutive time nodes is obtained. For example, on April 13th, there were 15 orders; on April 14th, there were 20 orders; on April 15th, there were 13 orders, and so on. Within a continuous time node, it is possible that at a specific time node, the package volume surges, for example, to 100,000 orders. Therefore, when discretizing features, to ensure that the processed data corresponds to a smaller data space, different isometry parameters are used for different order volumes. That is, the larger the package volume, the larger the isometry parameter.

[0118] Specifically, an initial dispersion coefficient is set, and the package volumes corresponding to each time node in the fourth time series sub-data are compared to determine the minimum package volume, which is used as the benchmark value. The planned package volume corresponding to the time node in the fourth time series sub-data is obtained. Since the resource acquisition space within the time series bucket is known, and the resource acquisition space has its own planned package volume for each time node, the planned package volume corresponding to the time node in the fourth time series sub-data can be obtained.

[0119] For the fourth time series sub-data, the planned package volume corresponding to each time node is compared with the baseline value. For example, the multiple relationship between the planned data and the baseline value is determined, and the modified dispersion coefficient is determined based on this. For example, if the ratio of the planned data to the baseline value is between 2 and 3, the modified dispersion coefficient is determined to be 2; if the ratio of the planned data to the baseline value is between 3 and 4, the modified dispersion coefficient is determined to be 3.

[0120] The baseline value is modified using the initial and revised discrete coefficients. For example, the product of the initial, revised, and baseline values is calculated and used as the equidistant parameter for that time node. After obtaining the equidistant parameter, the package volume is processed using the equidistant parameter for that time node to obtain the package volume at that time node after discretization. Similarly, the second time series sub-data for the time series bucket can be obtained.

[0121] Different modified discrete coefficients are used for different package quantities, that is, the overfitting of the model can be prevented by using a step-by-step discretization method.

[0122] Step S303: Use the package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time.

[0123] The package volume prediction model is configured to output a package volume prediction result based on the second time series data and the package plan data at the first prediction time. Figure 2 Step S203 of the illustrated embodiment is not limited herein.

[0124] The package volume prediction method provided in this embodiment further refines the regional dimension to the resource acquisition space dimension, which can improve data diversity. At the same time, the data in the time series bucket is processed according to the life cycle of the time series bucket to remove abnormal data. On the one hand, it reduces the data processing volume, and on the other hand, it ensures the reliability of the second time series data.

[0125] In this embodiment, a package volume prediction method is provided, which can be used in logistics scenarios. Figure 4 is a flowchart of a method for predicting package volume according to an embodiment of the present disclosure. Figure 4 As shown, the process includes the following steps:

[0126] Step S401: Acquire first time series data within a first historical period and package planning data at a first predicted time.

[0127] The first time series data is used to indicate the relationship between the time node and the package information. The package information includes the actual package quantity and the regional information. The regional information is used to indicate the region to which the resource acquisition space corresponding to the package belongs. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0128] Step S402: preprocess the first time series data to obtain second time series data.

[0129] The preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization. Figure 2 Step S202 of the embodiment shown, or, see Figure 3 Step S302 of the illustrated embodiment will not be described in detail here.

[0130] Step S403: Use the package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time.

[0131] The package volume prediction model is configured to output a package volume prediction result based on the second time series data and the package planning data of the first prediction time.

[0132] Exemplarily, the above step S403 includes:

[0133] Step S4031 , corresponding to the time series bucket, obtains prediction prompt information based on the second time series sub-data of the time series bucket and the package plan data of the first prediction time.

[0134] The second time series sub-data represents historical time series data, while the package plan data at the first forecast time guarantees future plan data and time. Prediction prompt information is generated based on the historical time series data, future plan data, and future time. For example, the prediction prompt information can be generated based on a prompt template, which includes fixed fields and corresponding variables. The second time series sub-data and the package plan data at the first forecast time are used as variables and filled into the variables of the corresponding fixed fields to generate the prediction prompt information.

[0135] In some optional embodiments, the prediction prompt information includes input prompt information and output prompt information. The input prompt information includes an input identifier, package plan data for a first prediction time, the first prediction time, and package target data, actual package data, and historical time obtained based on the second time series sub-data of the time series bucket; the output prompt information includes a predicted package quantity field and a predicted duration.

[0136] Specifically, if you want to predict the package volume for N days, then in the input prompt information, there are N pairs of package plan data for the first prediction time and the first prediction time, and each data pair corresponds one-to-one with the first prediction time; if the first historical duration obtained is M days, there are M groups of package target data, actual package data and historical time, and each group of data corresponds one-to-one with the historical time.

[0137] The output prompt information includes a predicted package quantity field and a predicted duration. The predicted duration corresponds to the first predicted time in the input prompt information, for example, N days.

[0138] The prediction prompt information has input prompt information and output prompt information, that is, input-output data pairs. Since the time series data of the logistics package volume forecast contains both future known covariates and historical multivariate information, the historical package volume and planned data are included in the prediction prompt information in the form of input-output data pairs, so that the package volume forecast model can capture the correlation between these variables, ensuring the forecast accuracy of the package volume forecast model.

[0139] Step S4032: input the prediction prompt information into the package quantity prediction model to obtain the first package prediction quantity of the time series bucket corresponding to the first prediction time.

[0140] After obtaining the forecast prompt information, it is input into the package volume prediction model to obtain the first package volume forecast. Since the forecast prompt information is obtained based on the second time series sub-data, and the second time series sub-data corresponds to a time series bucket, the obtained first package volume forecast also corresponds to the time series bucket. For example, the forecast is for the daily package volume in the time series bucket over the next N days.

[0141] Step S4033: The first package prediction quantity of the time series buckets corresponding to the same area at the first prediction time is merged to obtain the package quantity prediction result of the same area at the first prediction time.

[0142] Because resource acquisition space in the same region may be distributed across multiple time series buckets, the mapping relationship between region, time series bucket, and resource acquisition space can be used to query all time series buckets belonging to the same region. The predicted first package volume for the queried time series buckets at the same prediction time node is accumulated to obtain the package volume forecast for the same region at the first prediction time node.

[0143] The package volume prediction method provided in this embodiment uses a package volume prediction model that outputs the first package volume prediction for each time series bucket at the first prediction time. This output is then combined with the first package volume predictions for time series buckets corresponding to the same region to obtain the package volume prediction result for the same region at the first prediction time. The package volume prediction model is constructed based on a large language model and uses probability value sampling to make predictions. This allows for robustness despite some fluctuation in the predicted values generated for a single bucket.

[0144] In some optional implementations, a training process of a package volume prediction model is further included, specifically including:

[0145] Step d1: Configure the package volume prediction model.

[0146] Step d2: Create a sample data set, where the sample data set includes the third time series data within the second historical time period.

[0147] In step d3, a package quantity prediction model is trained using the sample data set. The package quantity prediction model is trained to output a predicted package quantity at a specified time.

[0148] The package volume prediction model can be built on a large language model base, for example, a non-manually aligned base model. Training the package volume prediction model is based on a sample dataset, so a sample dataset must be created before training.

[0149] The sample data set includes the third time series data within the second historical period. The generation method of the third time series data is similar to that of the second time series data. For details, see Figure 3 The generation process of the second time series data in the illustrated embodiment will not be described in detail here.

[0150] For example, the generation process of the first time series data may differ from that of the second time series data, including the following: the first time series data is obtained based on the time series data of the first historical duration, and the third time series data is obtained based on the time series data of the third historical duration. For example, the first historical duration is 14 days, the third historical duration is 300 days, and so on, but this is not limited to the above examples and other durations are also possible. It should be noted that the third historical duration is longer than the first historical duration.

[0151] For example, Figure 5 The process of model training and model inference is shown. Figure 5 The solid line in the figure represents the model training process, and the dotted line represents the model inference process.

[0152] Taking the third historical period of 300 days as an example, the training process typically requires tens of thousands of time series data points for model training, as there are at least hundreds of thousands of trainable parameters to effectively prevent overfitting during training. Especially during business growth phases, future package volumes often exceed past ones. For time series forecasting tasks, the model's training data should include possible future time series segments for more accurate predictions.

[0153] Therefore, if only regional training data is used, the number of samples over the past 300 days may be only a few hundred to a few thousand, and the historical package volume characteristics in the training data are generally lower than the predicted package volume characteristics. Therefore, during data engineering, it is necessary to further refine regional predictions to a finer granularity, such as down to the resource acquisition space dimension. This approach can significantly increase data diversity, better meet model training requirements, and enhance model generalization and prediction accuracy.

[0154] Since each resource acquisition space generates hundreds of sample data points, and thousands of resource acquisition spaces generate hundreds of thousands of sample data points, this can greatly increase data diversity. However, the daily package volume varies significantly between different resource acquisition spaces. While the resource acquisition space dimension increases the diversity of training data, it can also reduce the stability of model training.

[0155] In order to balance the diversity of samples and the stability of the target variable, the acquired third time series data can be preprocessed first, including but not limited to: bucketing, determining the life cycle, and discretizing features.

[0156] Specifically, we first set the number of time series buckets, then use a hashing method to bucket the resource acquisition space. After bucketing, we obtain the sum of the historical package volume of the resource acquisition space at the same time point within each time series bucket as the historical package volume feature of the time series bucket at that time point. This method improves data diversity while maintaining the stability of the target variable, thereby improving model training results.

[0157] Considering that some resource acquisition spaces may have operated for only a few months before being shut down in the past 300 days, we define a time series bucket lifecycle, with the first package quantity in a time series bucket being the start of the lifecycle and the last package quantity in the life series bucket being the end of the life cycle. If there is no package quantity data for a particular day during the life cycle of a time series bucket, zeros are added to the data, and the time series data is truncated at the end of the bucket life cycle, retaining only the time series data for each time series bucket from the beginning to the end of its life cycle.

[0158] For example, if the third time series data is data from the past 300 days, and for the resource acquisition space within the time series bucket, no data was generated in the first 30 days, then the number of days of valid time series data is considered to be 270 days. Then, the data for these 270 days is subsequently supplemented and divided into time series segments to obtain time series data for the first historical duration. For example, if 14 days of historical data is used to predict the package volume for the next 4 days, the length of the time series segment is 18 days; if 14 days of historical data is used to predict the package volume for the next 2 days, the length of the time series segment is 16 days.

[0159] The sensitivity of the package volume forecasting model to recent trends in time series data can be controlled by setting the length of each training sample's historical time series data. Shorter historical series lengths increase sensitivity to trend changes, while longer historical series lengths lead to more stable forecasts.

[0160] Because large language models excel at processing text and discrete data, continuous multivariate features need to be discretized when using them for time series regression. For example, for known future planned data, equal frequency discretization, equal distance discretization, or step discretization can be used. Step discretization is recommended to prevent overfitting.

[0161] Stepwise discretization uses the minimum value of the package volume in the time series data as the baseline, and determines the initial equidistance by multiplying the baseline by the initial dispersion coefficient (coeff, initialized here to 0.7). Specifically, let multiple = planned data (planned_data) / baseline value. The larger the multiple, the larger the modified dispersion coefficient. For example, for the part where the planned data is less than the baseline value, the equidistance value is the baseline value multiplied by the coefficient, that is, equidistance = baseline * coeff; for the part where the planned data is greater than the baseline value and less than or equal to 2 times the baseline value, the modified dispersion coefficient is 2, equidistance = baseline * coeff * 2; for the part where the planned data is greater than 2 times the baseline value and less than or equal to 3 times the baseline value, the modified dispersion coefficient is 3, equidistance = baseline * coeff * 3, and so on.

[0162] Since the time series data for logistics package volume forecasting contains both future known covariates and historical multivariate information, in order to capture the correlation between these variables, a large language model can be used for supervised fine-tuning (SFT) to learn the implicit pattern information in a large amount of training data, thereby making generative predictions.

[0163] In order to enable the large language model to learn future known covariates and historical multivariate information in time series data, input-output data pairs can be constructed as training data.

[0164] Exemplarily, the sample data set includes multiple sample subsets, and the sample subsets include sample time series data of a first duration, where the sample time series data of the first duration is obtained by sampling time series data of a second duration, and the second duration is greater than the first duration.

[0165] After dividing the time series data into time series segments, the obtained time series data is the second time series data. Based on this, sampling can be performed to obtain sample time series data of the first time series data. For example, if the second time series data is 14 items and the first time series data is 7 days, the time series data of the past 14 days can be randomly sampled to obtain 7 days of time series data, which can be used as the sample time series data.

[0166] By sampling the time series data of the second time length to obtain the time series data of the first time length as a sample subset, overfitting can be prevented and the robustness of the training model can be increased.

[0167] Exemplarily, the sample data set includes a first sample subset and a second sample subset. The first sample subset contains the package volume at a specified time node, and the second sample subset is obtained by performing sample enhancement on the first sample subset.

[0168] The training data may contain very important data, such as data at designated time points. At designated time points, package volume increases significantly, but the amount of data in this area is relatively small, typically only a few days a month. For these reasons, we need to randomly sample the data that includes designated time points in the output data of the training samples. This ultimately generates a sample size for designated time points, ensuring a near 1:1 ratio between the sample size of non-designated time points and the sample size of the designated time points, totaling hundreds of thousands of samples.

[0169] There may be a sharp increase in the number of packages at a specified time node, but the sample size at the specified time node is relatively small. Therefore, by performing sample enhancement on the first sample subset corresponding to the specified time node and increasing the sample size at the specified time node, it is ensured that the model training results have a high accuracy in predicting the package volume at the specified time node.

[0170] After obtaining the training dataset, the package volume prediction model is trained. For example, since the input and output of the large language model are both textual data, while time series prediction is typically purely numerical data, the large language model needs to be fine-tuned to ensure that the model's input and output are both numerical and that it memorizes the patterns of the training data. During fine-tuning, eight 32GB v100 graphics cards can be used for parallel training, and video memory usage can be reduced by reducing the training batch_size.

[0171] When selecting the base model for the package volume prediction model, a non-manually aligned base model was chosen to ensure that the output after fine-tuning the model returned to numeric values as much as possible. To conserve video memory, a large model of approximately 6 billion (6 Bytes) was selected.

[0172] The fine-tuning method uses Low Rank Adaptation (LORA) technology, meaning that LORA is used to fine-tune the model without adding additional time to model inference. Furthermore, when reasoning with the package volume prediction model, prediction results are output using probability-based sampling. This allows for some fluctuation in the prediction values generated for a single time series bucket, but the final prediction results for each region, derived through time series bucket-to-region aggregation, demonstrate robustness.

[0173] As a specific application example of the embodiment of the present disclosure, in the logistics scenario, the logistics party Figure 1 The terminal 101 shown submits a package quantity prediction request, which includes a prediction area and a prediction time. Figure 1After receiving the package quantity prediction request, the server 102 executes the package quantity prediction method described in the embodiment of the present disclosure to obtain the predicted package quantity for the predicted area within the predicted time period, and feeds the prediction result back to the terminal 101. Subsequently, the logistics party performs logistics scheduling based on the predicted package quantity.

[0174] This embodiment also provides a package volume prediction device for implementing the aforementioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.

[0175] This embodiment provides a package quantity prediction device, such as Figure 6 Shown, including:

[0176] Acquisition module 601 is used to obtain first time series data within a first historical period and package plan data for a first predicted time. The first time series data is used to indicate the relationship between time and package information. The package information includes actual package quantity and regional information. The regional information is used to indicate the region to which the resource acquisition space corresponding to the package belongs.

[0177] The preprocessing module 602 is used to preprocess the first time series data to obtain second time series data, where the preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization.

[0178] The prediction module 603 is used to use the package volume prediction model to predict the package volume at the first prediction time to obtain the package volume prediction result of the area at the first prediction time. The package volume prediction model is configured to output the package volume prediction result based on the second time series data and the package plan data at the first prediction time.

[0179] In some optional implementations, the pre-processing module 602 includes:

[0180] The bucketing unit is configured to bucket the first time series data based on the resource acquisition space in the region information to obtain data in the time series bucket, where the data in the time series bucket includes time series sub-data of multiple resource acquisition spaces in the same region.

[0181] The lifecycle determination unit is used to determine the lifecycle of a time series bucket based on the data in the time series bucket.

[0182] The data processing unit is configured to process the data in the time series bucket based on the life cycle to obtain the second time series sub-data of the time series bucket, where the second time series data includes the second time series sub-data of all the time series buckets.

[0183] In some optional embodiments, the bucketing unit includes:

[0184] The first determining subunit is configured to determine a resource acquisition space within a same region in the first time series data and the number of time series buckets.

[0185] The bucketing sub-unit is configured to bucket the first time series data based on the resource acquisition space and the number of time series buckets in the same region to obtain the data in the time series buckets.

[0186] In some optional implementations, the lifecycle determination unit includes:

[0187] A time acquisition subunit is used to acquire, for the data in the time series bucket, a first time when the package quantity first exists and a second time when the package quantity last exists;

[0188] The life cycle determination subunit is configured to use the first time as the start time of the life cycle and the second time as the end time of the life cycle.

[0189] In some optional embodiments, the data processing unit includes:

[0190] A filling subunit fills the package quantity of the entire life cycle with the time series sub-data of the resource acquisition space in the time series bucket to obtain third time series sub-data of the resource acquisition space;

[0191] The merging sub-unit is configured to merge the package quantities at the same time node for the third time series sub-data of all the resource acquisition spaces in the time series bucket to obtain the second time series sub-data of the time series bucket.

[0192] In some optional embodiments, the merging subunit includes:

[0193] The accumulation sub-unit is used to accumulate the package quantity at the same time node for all the third time series sub-data in the time series bucket to obtain the fourth time series sub-data.

[0194] The discretization subunit is configured to perform feature discretization processing on the fourth time series sub-data to obtain second time series sub-data.

[0195] In some optional embodiments, the discretization subunit includes:

[0196] The reference value determination subunit is used to determine the minimum value of the initial discrete coefficient and the package quantity in the fourth time series sub-data, and use the minimum value as the reference value.

[0197] The correction coefficient determination subunit is used to determine the correction dispersion coefficient based on the relationship between the planned package quantity corresponding to the time node in the fourth time series sub-data and the reference value.

[0198] The correction subunit is used to correct the reference value based on the initial discrete coefficient and the corrected discrete coefficient to obtain the equidistant parameters of the time node.

[0199] The package quantity processing subunit is used to process the package quantity corresponding to the time node based on the equidistant parameter to obtain the second time series sub-data.

[0200] In some optional embodiments, the prediction module includes:

[0201] The prompt information determining unit is configured to obtain prediction prompt information corresponding to the time series bucket based on the second time series sub-data of the time series bucket and the package plan data at the first prediction time.

[0202] The prediction unit is used to input the prediction prompt information into the package quantity prediction model to obtain the first package prediction quantity of the time series bucket corresponding to the first prediction time.

[0203] The fusion unit is used to fuse the first package prediction quantity of the time series buckets corresponding to the same area at the first prediction time to obtain the package quantity prediction result of the same area at the first prediction time.

[0204] In some optional implementations, the predicted prompt information includes input prompt information and output prompt information;

[0205] The input prompt information includes an input identifier, package plan data for the first prediction time, the first prediction time, and package target data obtained based on the second time series sub-data of the time series bucket, actual package data, and historical time;

[0206] The output prompt information includes the predicted package quantity field and the predicted duration.

[0207] In some optional implementations, the package quantity prediction device further includes:

[0208] Configuration module, used to configure the package volume prediction model.

[0209] The creation module is used to create a sample data set, where the sample data set includes third time series data of the second package within the second historical time period.

[0210] The training module is used to train a package volume prediction model using a sample data set, and the package volume prediction model is trained to output a predicted package volume at a specified time.

[0211] In some optional embodiments, the sample data set includes multiple sample subsets, and the sample subsets include sample time series data of a first duration, where the sample time series data of the first duration is obtained by sampling time series data of a second duration, and the second duration is greater than the first duration.

[0212] In some optional implementations, the sample data set includes a first sample subset and a second sample subset, the first sample subset includes the package volume at a specified time node, and the second sample subset is obtained by performing sample enhancement on the first sample subset.

[0213] The package volume prediction device provided by the embodiments of the present disclosure can execute the package volume prediction method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. For example, when using the package volume prediction model to predict the package volume, not only the actual historical package volume is taken into account, but also the planned data at the time of prediction is taken into account. At the same time, the prediction results can be quickly adjusted according to recent data changes, thereby improving the accuracy of the package volume prediction, thereby meeting the efficient production capacity arrangement of the logistics business. The further functional description of each of the above modules and units is the same as that of the corresponding embodiment above, and will not be repeated here.

[0214] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0215] The following specific reference Figure 7 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a memory 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device are also stored in the RAM 703. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0216] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0217] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from memory 708, or installed from ROM 702. When executed by processor 701, the computer program performs the aforementioned functions defined in the package volume prediction method of the embodiments of the present disclosure.

[0218] Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0219] The present disclosure also provides a computer-readable storage medium. The method according to the embodiment of the present disclosure can be implemented in hardware, firmware, or as a computer code that can be recorded on a storage medium, or downloaded over a network and originally stored in a remote storage medium or a non-transitory machine-readable storage medium and then stored in a local storage medium. Thus, the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the package volume prediction method shown in the above embodiment is implemented.

[0220] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0221] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for predicting package volume, characterized in that: The method comprises: Obtaining first time series data within a first historical time period and package plan data for a first predicted time period, wherein the first time series data is used to indicate a relationship between a time node and package information, the package information includes an actual package quantity and region information, and the region information is used to indicate a region to which a resource acquisition space corresponding to the package belongs; Preprocessing the first time series data to obtain second time series data, wherein the preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization; The package volume prediction model is used to predict the package volume at the first prediction time to obtain the package volume prediction result of the area at the first prediction time. The package volume prediction model is configured to output the package volume prediction result based on the second time series data and the package plan data at the first prediction time.

2. The method according to claim 1, characterized in that The preprocessing of the first time series data to obtain second time series data includes: Bucketing the first time series data based on the resource acquisition space in the region information to obtain data in the time series bucket, where the data in the time series bucket includes time series sub-data of multiple resource acquisition spaces in the same region; Determining a life cycle of the time series bucket based on the data in the time series bucket; The data in the time series bucket is processed based on the life cycle to obtain second time series sub-data of the time series bucket, where the second time series data includes the second time series sub-data of all the time series buckets.

3. The method according to claim 2, characterized in that The step of bucketing the first time series data based on the resource acquisition space in the region information to obtain data in the time series buckets includes: Determining a resource acquisition space within the same region in the first time series data and the number of the time series buckets; Based on the resource acquisition space in the same region and the number of the time series buckets, the first time series data is bucketed to obtain data in the time series buckets.

4. The method according to claim 2, characterized in that The determining the life cycle of the time series bucket based on the data in the time series bucket includes: For the data in the time series bucket, obtain the first time when the package quantity first exists and the second time when the package quantity last exists; The first time is used as the start time of the life cycle, and the second time is used as the end time of the life cycle.

5. The method according to claim 2, characterized in that The processing of the data in the time series bucket based on the life cycle to obtain second time series sub-data of the time series bucket includes: Completing the package quantity of the entire life cycle for the time series sub-data of the resource acquisition space in the time series bucket to obtain third time series sub-data of the resource acquisition space; For the third time series sub-data of all the resource acquisition spaces in the time series bucket, the package quantities at the same time node are merged to obtain the second time series sub-data of the time series bucket.

6. The method according to claim 5, characterized in that The third time series sub-data of all resource acquisition spaces in the time series bucket are merged with the package quantities at the same time node to obtain the second time series sub-data of the time series bucket, including: For all the third time series sub-data in the time series bucket, the package quantity at the same time node is accumulated to obtain fourth time series sub-data; Feature discretization processing is performed on the fourth time series sub-data to obtain the second time series sub-data.

7. The method according to claim 6, characterized in that The performing feature discretization processing on the fourth time series sub-data to obtain the second time series sub-data includes: Determine an initial dispersion coefficient and a minimum value of the package quantity in the fourth time series sub-data, and use the minimum value as a reference value; Determine a modified dispersion coefficient based on the relationship between the planned package volume corresponding to the time node in the fourth time series sub-data and the benchmark value; Correcting the reference value based on the initial discrete coefficient and the corrected discrete coefficient to obtain the equidistant parameter of the time node; The package quantity corresponding to the time node is processed based on the equidistant parameter to obtain the second time series sub-data.

8. The method according to claim 2, characterized in that The method of using the package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time includes: Corresponding to the time series bucket, obtaining prediction prompt information based on the second time series sub-data of the time series bucket and the package plan data of the first prediction time; Inputting the prediction prompt information into the package quantity prediction model to obtain the first package prediction quantity of the time series bucket corresponding to the first prediction time; The first package prediction quantities of the time series buckets corresponding to the same area at the first prediction time are merged to obtain the package quantity prediction result of the same area at the first prediction time.

9. The method according to claim 8, characterized in that The prediction prompt information includes input prompt information and output prompt information; The input prompt information includes an input identifier, package plan data for the first prediction time, the first prediction time, and package target data, package actual data, and historical time obtained based on the second time series sub-data of the time series bucket; The output prompt information includes a predicted package quantity field and a predicted duration.

10. The method according to claim 1, characterized in that The method further comprises: Configure the package volume prediction model; Creating a sample data set, where the sample data set includes third time series data within a second historical period; The package quantity prediction model is trained using the sample data set, and the package quantity prediction model is trained to output a predicted package quantity at a specified time.

11. The method according to claim 10, characterized in that The sample data set includes multiple sample subsets, each of which includes sample time series data of a first duration, where the sample time series data of the first duration is obtained by sampling time series data of a second duration, and the second duration is greater than the first duration.

12. The method according to claim 10, characterized in that The sample data set includes a first sample subset and a second sample subset. The first sample subset contains the package volume at a specified time node, and the second sample subset is obtained by performing sample enhancement on the first sample subset.

13. A package quantity prediction device, characterized in that: The device comprises: an acquisition module, configured to acquire first time series data within a first historical time period and package plan data for a first predicted time period, wherein the first time series data indicates a relationship between time and package information, the package information includes an actual package quantity and region information, and the region information indicates a region to which a resource acquisition space corresponding to the package belongs; a preprocessing module, configured to preprocess the first time series data to obtain second time series data, wherein the preprocessing includes at least one of time series data segmentation, time series data organization, and time series data discretization; A prediction module is used to use a package volume prediction model to predict the package volume at the first prediction time to obtain a package volume prediction result for the area at the first prediction time, and the package volume prediction model is configured to output the package volume prediction result based on the second time series data and the package plan data at the first prediction time.

14. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the package quantity prediction method according to any one of claims 1 to 12 by executing the computer instructions.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the package quantity prediction method according to any one of claims 1 to 12.

16. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the package quantity prediction method according to any one of claims 1 to 12.