Data processing method and device, electronic equipment and storage medium

By calculating the nearest neighbor relationship, type similarity, and trend similarity between the target date and historical dates, highly relevant historical dates are selected for training. This solves the problem of insufficient recognition accuracy of pre-trained models in the package processing system and achieves more efficient data processing and model training results.

CN120930841APending Publication Date: 2025-11-11SF TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410578435.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies suffer from low recognition accuracy when using pre-trained models for data processing, especially in package processing systems, where the failure to effectively select historical dates with high similarity to future dates for training leads to insufficient model recognition accuracy.

Method used

By calculating the nearest neighbor relationship, date type similarity, and trend similarity between the target date and historical dates, multiple target historical dates are selected. Based on this data, a package processing volume prediction model is trained, and the date selection process is optimized using weighted calculation and date decay function.

Benefits of technology

It improves the accuracy and generalization ability of the parcel processing volume prediction model, reduces the impact of invalid data, lowers computational complexity, and improves the model's recognition accuracy and training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930841A_ABST
    Figure CN120930841A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the steps of obtaining a target date of a to-be-predicted parcel handling capacity and historical parcel handling capacity data within a preset time range; calculating the relevance between each historical date and the target date based on the neighbor relation between the target date and each historical date in the preset time range; obtaining first date type data of the target date and second date type data of each historical date, and calculating date type similarity; obtaining first parcel handling capacity trend data corresponding to the target date and second parcel handling capacity trend data corresponding to each historical date, and calculating trend similarity; and determining a plurality of target historical dates based on the relevance, the date type similarity and the trend similarity, and training a parcel handling capacity prediction model. According to the embodiment of the invention, the accuracy of data processing can be improved, so that the recognition accuracy of the pre-training model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a data processing method and apparatus, electronic device and storage medium. Background Technology

[0002] Data processing refers to a series of operations including data collection, organization, storage, analysis, and presentation. For resource application scenarios with large amounts of data, training a pre-trained model for the system using this massive amount of data can lead to high computational costs, reducing the model's accuracy. For example, in a parcel processing system, "resources" refers to the volume of parcels processed. A pre-trained model for a resource application scenario can predict resource availability for future dates, allowing for advance planning.

[0003] Therefore, existing technologies typically combine data processing with data partitioning. This involves using historical resource quantities to predict future resource quantities and then selecting appropriate training data from a large dataset based on the prediction results to improve the accuracy of the pre-trained model for the resource application scenario. However, this method of selecting training data based on the most recent data of the same date type results in low data similarity, leading to low accuracy in the pre-trained model for resource application scenarios. Therefore, improving the accuracy of data processing—specifically, selecting historical dates with high similarity for future dates and then using this selected historical date data to train the pre-trained model (such as a package processing volume prediction model) for the resource application scenario system—has become a pressing technical problem. Summary of the Invention

[0004] The main objective of this application is to propose a data processing method, apparatus, electronic device, and storage medium that can improve the accuracy of data processing. Specifically, it selects historical dates with high similarity to future dates, and then trains a pre-trained model (such as a parcel processing volume prediction model) corresponding to the system of the resource application scenario based on the data of the selected historical dates, so as to improve the recognition accuracy of the pre-trained model.

[0005] To achieve the above objectives, a first aspect of this application provides a data processing method, the method comprising:

[0006] Obtain the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range;

[0007] Based on the nearest neighbor relationship between the target date and each historical date within the preset time range, the correlation between each historical date and the target date is calculated;

[0008] Obtain first date type data for the target date and second date type data for each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data;

[0009] Obtain the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data;

[0010] Based on the correlation, the date type similarity, and the trend similarity, multiple target historical dates are determined within the preset time range, and a parcel processing volume prediction model is trained based on the historical parcel processing volume data corresponding to the multiple target historical dates.

[0011] In some embodiments, determining multiple target historical dates within the preset time range based on the correlation, the date type similarity, and the trend similarity includes:

[0012] Obtain a pre-set first weight, second weight, and third weight; the first weight corresponds to the correlation between each historical date and the target date, the second weight corresponds to the date type similarity between each historical date and the target date, and the third weight corresponds to the trend similarity between each historical date and the target date;

[0013] The correlation is weighted based on the first weight to obtain the first weighted data;

[0014] The date type similarity is weighted based on the second weight to obtain the second weighted data;

[0015] The trend similarity is weighted based on the third weight to obtain the third weighted data.

[0016] Based on the first weighted data, the second weighted data, and the third weighted data, a date selection is performed within the preset time range to determine multiple target historical dates.

[0017] In some embodiments, the step of selecting dates within the preset time range based on the first weighted data, the second weighted data, and the third weighted data to determine multiple target historical dates includes:

[0018] For each historical date, the first weighted data, the second weighted data, and the third weighted data are summed to obtain the candidate similarity for each historical date;

[0019] The candidate similarities of multiple historical dates within the preset time range are sorted in descending order to obtain the target historical date sequence;

[0020] Select a preset date threshold number of target historical dates from the target historical date sequence.

[0021] In some embodiments, calculating the correlation between each historical date and the target date based on the nearest neighbor relationship between the target date and each historical date within the preset time range includes:

[0022] Obtain the preset offset reference data and date decay parameters;

[0023] The historical date offset value is determined based on the offset reference data and the historical date;

[0024] Determine the target date offset value based on the historical date and the target date;

[0025] Based on a preset date decay function and the date decay parameters, the historical date offset value is decayed to obtain a historical decay coefficient;

[0026] Based on the date decay function and the date decay parameter, the target date offset value is decayed to obtain the target decay coefficient;

[0027] The correlation is determined by comparing the historical attenuation coefficient and the target attenuation coefficient.

[0028] In some embodiments, the date decay parameter includes a day decay sub-parameter and a week decay sub-parameter. The step of performing date decay on the historical date offset value based on a preset date decay function and the date decay parameter to obtain a historical decay coefficient includes:

[0029] The difference between the day decay sub-parameter and the historical date offset value is calculated to obtain the historical day decay difference.

[0030] The historical week offset value is determined based on the historical date offset value;

[0031] The difference between the cycle attenuation sub-parameter and the historical cycle offset value is calculated to obtain the historical cycle attenuation difference.

[0032] The historical attenuation coefficient is obtained by calculating the attenuation function based on the historical daily attenuation difference and the historical weekly attenuation difference.

[0033] In some embodiments, the step of performing date decay on the target date offset value based on the date decay function and the date decay parameter to obtain the target decay coefficient includes:

[0034] The difference between the day attenuation sub-parameter and the target date offset value is calculated to obtain the target day attenuation difference;

[0035] Determine the target week offset value based on the target date offset value;

[0036] The difference between the cycle attenuation sub-parameter and the target cycle offset value is calculated to obtain the target cycle attenuation difference;

[0037] The target attenuation coefficient is obtained by calculating the attenuation function based on the date attenuation function to determine the target daily attenuation difference and the target weekly attenuation difference.

[0038] In some embodiments, obtaining first date type data for the target date and second date type data for each of the historical dates includes:

[0039] Extract the date type from the target date to obtain the target date type;

[0040] For each historical date, extract the date type to obtain the historical date type for each historical date;

[0041] The target date type is encoded based on a preset date type encoding model to obtain the first date type data;

[0042] Based on the date type encoding model, the historical date type of each historical date is encoded to obtain the second date type data of each historical date.

[0043] In some embodiments, obtaining the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each of the historical dates includes:

[0044] Obtain the preset trend prediction model;

[0045] Based on the trend prediction model, the trend data of the first package processing volume and the target date are predicted to obtain the target trend data;

[0046] Based on the trend prediction model, the trend data of the second parcel processing volume and the historical date are used to predict the trend, and historical trend data is obtained.

[0047] In some embodiments, the trend prediction model includes a potential energy prediction sub-model and an energy level prediction sub-model; the step of predicting the target trend data based on the trend prediction model of the first package processing volume trend data and the target date includes:

[0048] Based on the potential energy prediction sub-model, the first package processing volume trend data and the target date are used to predict the potential energy to obtain the target potential energy data.

[0049] Based on the energy level prediction sub-model, energy level prediction is performed on the first package processing volume trend data and the target date to obtain target energy level data;

[0050] The target trend data is constructed based on the target potential energy data and the target energy level data.

[0051] In some embodiments, the step of performing trend prediction on the second parcel processing volume trend data and the historical date based on the trend prediction model to obtain historical trend data includes:

[0052] Based on the potential energy prediction sub-model, the potential energy is predicted using the trend data of the second package processing volume and the historical date to obtain historical potential energy data.

[0053] Based on the energy level prediction sub-model, energy level prediction is performed on the second package processing volume trend data and the historical date to obtain historical energy level data;

[0054] The historical trend data is constructed based on the historical potential energy data and the historical energy level data.

[0055] To achieve the above objectives, a second aspect of this application provides a data processing apparatus, the apparatus comprising:

[0056] The data acquisition module is used to acquire the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range;

[0057] The correlation calculation module is used to calculate the correlation between each historical date and the target date based on the nearest neighbor relationship between the target date and each historical date within the preset time range;

[0058] A type calculation module is used to obtain first date type data of the target date and second date type data of each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data;

[0059] The trend prediction module is used to acquire the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data.

[0060] The date determination module is used to determine multiple target historical dates within the preset time range based on the correlation, the date type similarity, and the trend similarity, and to train a parcel processing volume prediction model based on the historical parcel processing volume data corresponding to the multiple target historical dates.

[0061] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0062] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0063] This application proposes a data processing method, apparatus, electronic device, and storage medium. By considering the proximity relationship, date type similarity, and trend similarity between the target date and historical dates, it can more accurately filter out multiple historical dates with high correlation to the target date. This effectively reduces the impact of invalid data on model training, improves computational efficiency, and thus enhances the accuracy of the parcel processing volume prediction model. This application trains the parcel processing volume prediction model using historical parcel processing volume data corresponding to multiple selected target historical dates, enabling the model to better adapt to different date types and trend changes, thereby improving the model's generalization ability. Therefore, this application can improve the accuracy of data processing by selecting historical dates with high similarity to future dates. The pre-trained model (such as the parcel processing volume prediction model) corresponding to the resource application scenario system is then trained based on the selected historical date data to improve the recognition accuracy of the pre-trained model. Attached Figure Description

[0064] Figure 1 This is a flowchart of a data processing method provided in an embodiment of this application;

[0065] Figure 2 yes Figure 1 A flowchart of step S102 in the process;

[0066] Figure 3 yes Figure 2A flowchart of step S204 in the process;

[0067] Figure 4 yes Figure 1 A flowchart of step S103 in the process;

[0068] Figure 5 yes Figure 1 A flowchart of step S104 in the process;

[0069] Figure 6 yes Figure 1 A flowchart of step S105 in the process;

[0070] Figure 7 yes Figure 6 A flowchart of step S605 in the process;

[0071] Figure 8 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application;

[0072] Figure 9 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0074] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0076] First, let's analyze some of the terms used in this application:

[0077] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0078] Data processing refers to a series of operations including data collection, organization, storage, analysis, and presentation. For resource application scenarios with large amounts of data, training a pre-trained model for the system using this massive amount of data can lead to high computational costs, reducing the model's accuracy. For example, in a parcel processing system, "resources" refers to the volume of parcels processed. A pre-trained model for a resource application scenario can predict resource availability for future dates, allowing for advance planning.

[0079] Therefore, existing technologies typically combine data processing with data partitioning. This involves using historical resource quantities to predict future resource quantities and then selecting appropriate training data from a large dataset based on the prediction results to improve the accuracy of the pre-trained model for the resource application scenario. However, this method of selecting training data based on the most recent data of the same date type results in low data similarity, leading to low accuracy in the pre-trained model for resource application scenarios. Therefore, improving the accuracy of data processing—specifically, selecting historical dates with high similarity for future dates and then using this selected historical date data to train the pre-trained model (such as a package processing volume prediction model) for the resource application scenario system—has become a pressing technical problem.

[0080] Based on this, embodiments of this application provide a data processing method and apparatus, electronic device and storage medium, which can improve the accuracy of data processing, namely, selecting historical dates with high similarity to future dates, and then training a pre-trained model (such as a parcel processing volume prediction model) corresponding to the system of resource application scenarios based on the data of the selected historical dates, so as to improve the recognition accuracy of the pre-trained model.

[0081] This application provides a data processing method, apparatus, electronic device, and storage medium, which are specifically described through the following embodiments. First, the data processing method in the embodiments of this application is described.

[0082] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0083] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0084] The data processing method provided in this application relates to the field of artificial intelligence technology. The data processing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the data processing method, but is not limited to the above forms.

[0085] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0086] Figure 1 This is an optional flowchart of a data processing method provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.

[0087] Step S101: Obtain the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range;

[0088] Step S102: Based on the nearest neighbor relationship between the target date and each historical date within a preset time range, calculate the correlation between each historical date and the target date;

[0089] Step S103: Obtain the first date type data of the target date and the second date type data of each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data;

[0090] Step S104: Obtain the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data.

[0091] Step S105: Based on correlation, date type similarity and trend similarity, determine multiple target historical dates within a preset time range, and train a parcel processing volume prediction model based on the historical parcel processing volume data corresponding to the multiple target historical dates.

[0092] Steps S101 to S105 of this application embodiment, by considering the proximity relationship, date type similarity, and trend similarity between the target date and historical dates, can more accurately filter historical dates with high correlation to the target date, thereby improving the accuracy of the parcel processing volume prediction model, reducing the amount of data required for model training, and lowering computational complexity. This application also trains the parcel processing volume prediction model on historical parcel processing volume data corresponding to multiple target historical dates, enabling the model to better adapt to different date types and trend changes, thereby improving the model's generalization ability. Compared to related technologies that only select training data based on the most recent data within the same date type, this application embodiment considers three different dimensions: proximity relationship, date type similarity, and trend similarity, resulting in a strong correlation between the target historical dates and the target date. Therefore, using multiple target historical dates with stronger correlation as sample data for the parcel processing volume prediction model can improve the accuracy of data processing and the accuracy of the parcel processing volume prediction model.

[0093] In step S101 of some embodiments, package processing volume refers to package data processed within a certain period of time. For example, package processing volume can be the number of packages, the total weight of packages, etc., without specific limitations. In specific application scenarios, package processing volume can be used to measure the business volume of a logistics or express delivery company within a certain period, and can also be used to analyze logistics operational efficiency or predict future transportation demand. The target date is a specific future date on which the package processing volume needs to be predicted. Since the package processing volume for the predicted day is currently undetermined data, the target date can be the predicted day. The target date can be expressed as a combination of year, month, and day. For example, a specific target date can be set as June 18, 2024. The preset time range refers to a pre-defined period of time. Because the amount of historical date data is very large, and historical package processing volume data for historical dates with large intervals from the target date are not very useful for predicting the package processing volume for the target date, selecting the historical date to be judged according to the preset time range can improve data processing efficiency. The preset time range can be set according to actual needs. For example, it can be 180 (days), 360 (days), etc. Historical parcel processing volume data refers to the total number of parcels processed on each historical date. If the preset time range is 180 days, then the historical parcel processing volume data can be the daily parcel processing volume of logistics or courier companies over the past 180 days. Therefore, there are 180 historical parcel processing volume data points, each corresponding to the historical parcel processing volume data for the past 180 historical dates.

[0094] Please see Figure 2In some embodiments, step S102 may include, but is not limited to, steps S201 to S206:

[0095] Step S201: Obtain preset offset reference data and date decay parameters;

[0096] Step S202: Determine the historical date offset value based on the offset reference data and historical dates;

[0097] Step S203: Determine the target date offset value based on the historical date and the target date;

[0098] Step S204: Based on the preset date decay function and date decay parameters, the historical date offset value is subject to date decay to obtain the historical decay coefficient;

[0099] Step S205: Perform date decay on the target date offset value based on the date decay function and date decay parameters to obtain the target decay coefficient;

[0100] Step S206: Based on the correlation comparison between the historical attenuation coefficient and the target attenuation coefficient, determine the correlation.

[0101] In step S201 of some embodiments, the date decay parameter refers to a pre-set time interval parameter used for decay calculation. Setting the date decay parameter ensures the referenceability of historical dates; if the time interval between the target date and a historical date exceeds the set date decay parameter, the current historical date is considered unreliable. Offset reference data refers to reference data used to determine the offset value of historical dates. The offset reference data can be determined based on a preset time range; that is, the smallest historical date within the preset time range that is an interval from the target date can be used as the offset reference data. The offset reference data can also be any historical date selected within the preset time range; no specific limitation is made here.

[0102] It should be noted that, in order to simplify the calculation of dates, this application sets the total number of days in each month to 30 days in the subsequent embodiments.

[0103] In step S202 of some embodiments, the historical date refers to a specific date that occurred or has occurred in the past. Specifically, the historical date is expressed in the form of year, month, and day. For example, a specific historical date can be expressed as January 18, 2024. The historical date offset value refers to the date time interval offset value in days between the historical date and the offset reference data. Specifically, the historical date offset value can be determined using a date offset function, such as using the DATED IF function in Excel, etc., which is not limited here. For example, when the offset reference data is December 18, 2023, and the historical date is January 18, 2024, the corresponding historical date offset value is the interval between December 18, 2023, and January 18, 2024, which is 30 days.

[0104] In step S203 of some embodiments, the target date offset value may refer to the date time interval offset between the target date and the historical date. Furthermore, the target date offset value may be determined in days. For example, when the historical date is January 18, 2024, and the target date is June 18, 2024, the target date offset value is the interval between January 18, 2024, and June 18, 2024, which is 150 days.

[0105] In step S204 of some embodiments, the date decay parameter includes a daily decay sub-parameter and a weekly decay sub-parameter. The daily decay sub-parameter refers to the decay data of the time interval in days, and the weekly decay sub-parameter refers to the decay data of the time interval in weeks. Specifically, a daily decay sub-parameter can be preset, and the weekly decay sub-parameter can be the date decay data divided by 7 and then rounded to the nearest integer. Alternatively, a weekly decay sub-parameter can be preset, and the daily decay sub-parameter is the weekly decay data multiplied by 7. For example, the date decay parameters can be denoted as 180 (days) and 26 (weeks), where 180 (days) corresponds to the daily decay sub-parameter and 26 (weeks) corresponds to the weekly decay sub-parameter.

[0106] It should be noted that date decay refers to the process of substituting historical date offsets and date decay parameters into a date decay function for attenuation. A date decay function is a function that calculates the decay coefficient between two specific dates. The decay coefficient can be used to characterize the proximity between two specific dates. The closer a historical date is to the target date, the more reliable it is, and the larger the target decay coefficient calculated based on the date decay function will be. The date decay function used in this application can be a logarithmic function, a linear decay function, an exponential decay function, etc., and can be flexibly selected according to actual needs; no specific limitation is made here.

[0107] It should be noted that the historical decay coefficient refers to the coefficient calculated based on the date decay function and date decay parameters for the historical date offset value. The historical decay coefficient is used to characterize the proximity of the historical date and the offset reference data. Specifically, this embodiment utilizes the proximity relationship between the target date and the historical date to construct a decreasing date decay function to determine the target decay coefficient and the historical decay coefficient. The larger the target date offset value, the smaller the corresponding target decay coefficient; the larger the historical date offset value, the smaller the corresponding historical decay coefficient. In another embodiment, the date decay function conforms to the requirement of time-series effects progressing from near to far and can also be constructed based on the correlation coefficient between historical dates and target dates of the same type. Assuming that the correlation coefficient between historical dates and target dates of the same type decreases as the time interval offset increases, this decreasing linear fitting slope can be used as a reference line for fitting the date decay function. Specifically, the independent variable can be the time interval offset between the target date and the historical date, and the dependent variable can be the linear correlation coefficient between the target date and the historical date. The linear fitting slope of the function (i.e., the reference line for fitting the date decay function) can be calculated using historical data statistical methods.

[0108] This application's embodiments, by calculating the correlation between each historical date and the target date, can more accurately filter out historical dates with a high correlation to the target date, which helps improve the accuracy of subsequent parcel processing volume prediction. Secondly, by calculating the correlation between historical dates and the target date, the most relevant historical dates can be selected from multiple historical dates, thereby reducing the amount of data required for training the subsequent parcel processing volume prediction model and lowering computational complexity. Thirdly, the correlation calculation can also help to filter out useful historical data more quickly, shortening the training time of the parcel processing volume prediction model and improving its training efficiency. Finally, the correlation calculation results can also provide some interpretability for the parcel processing volume prediction model, helping to understand why certain historical dates have a greater impact on the prediction of the target date, thereby improving the interpretability of the parcel processing volume prediction model.

[0109] Please see Figure 3 In some embodiments, step S204 may include, but is not limited to, steps S301 to S304:

[0110] Step S301: Calculate the difference between the daily attenuation sub-parameter and the historical date offset value to obtain the historical daily attenuation difference;

[0111] Step S302: Determine the historical week offset value based on the historical date offset value;

[0112] Step S303: Calculate the difference between the cycle attenuation sub-parameter and the historical cycle offset value to obtain the historical cycle attenuation difference;

[0113] Step S304: Calculate the historical daily attenuation difference and historical weekly attenuation difference based on the date attenuation function to obtain the historical attenuation coefficient.

[0114] In step S301 of some embodiments, the historical day decay difference refers to the value obtained by subtracting the day decay sub-parameter from the historical date offset value. For example, the day decay sub-parameter can be 180 (days), the historical date offset value can be 30 (days), and the historical day decay difference can be 180-30=150 (days).

[0115] In step S302 of some embodiments, the historical week offset value refers to a value determined based on the number of weeks of the historical date offset value. Specifically, the historical week offset value is obtained by dividing the historical date offset value by 7 and then rounding it to the nearest integer. For example, if the historical date offset value can be 30 (days), then the historical week offset value is 30 ÷ 7 ≈ 4 (weeks).

[0116] In step S303 of some embodiments, the historical cycle attenuation difference refers to the value obtained by subtracting the cycle attenuation sub-parameter from the historical cycle offset value. For example, the cycle attenuation sub-parameter can be 26 (cycles), the historical cycle offset value can be 4 (cycles), and the historical cycle attenuation difference can be 26-4=22 (cycles).

[0117] In step S304 of some embodiments, the attenuation function calculation refers to the process of substituting the historical daily attenuation difference and the historical weekly attenuation difference into the date attenuation function and then performing attenuation calculation. For example, the date attenuation parameters can be denoted as 180 (days) and 26 (weeks), and the historical daily attenuation difference can be denoted as d. o The historical weekly decay difference can be denoted as ω. o The pre-defined date decay function can be L = lnd o +lnω o Then the historical attenuation coefficient is ln150+ln 22≈8.1.

[0118] In step S205 of some embodiments, the target attenuation coefficient refers to the attenuation coefficient obtained by processing the target date offset value according to the date attenuation function and the date attenuation parameter. That is, the target attenuation coefficient can be used to characterize the proximity of the target date and historical dates. Date attenuation refers to the process of attenuating the date attenuation parameter and the target date offset value by substituting them into the date attenuation function.

[0119] Specifically, step S205 includes: calculating the difference between the daily attenuation sub-parameter and the target date offset value to obtain the target daily attenuation difference; determining the target weekly offset value based on the target date offset value; calculating the difference between the weekly attenuation sub-parameter and the target weekly offset value to obtain the target weekly attenuation difference; and calculating the attenuation function based on the date attenuation function to obtain the target attenuation coefficient for the target daily attenuation difference and the target weekly attenuation difference.

[0120] It should be noted that the target daily attenuation difference is the value obtained by subtracting the target date offset from the daily attenuation sub-parameter. The target weekly offset is the value determined based on the number of weeks of the target date offset. The method for determining the target weekly offset is the same as the method for determining the historical weekly offset in step S302 above, and will not be repeated here. The target weekly attenuation difference is the value obtained by subtracting the target weekly offset from the weekly attenuation sub-parameter. Attenuation function calculation refers to the process of substituting the target daily attenuation difference and the target weekly attenuation difference into the date attenuation function to perform attenuation calculation. For example, if the daily attenuation sub-parameter can be 180 (days) and the target date offset can be 150 (days), then the target daily attenuation difference can be 180-150=30 (days). If the weekly attenuation sub-parameter can be 26 (weeks), the target weekly offset can be 150÷7≈21 (weeks), and the target weekly attenuation difference can be 26-21=5 (weeks). The target daily attenuation difference can be denoted as d. i The target periodic attenuation difference can be denoted as ω. i The pre-defined date decay function can be L = lnd i +lnω i In summary, the target attenuation coefficient can be ln30 + ln 5 ≈ 5.0.

[0121] In step S206 of some embodiments, correlation refers to the degree of correlation between the historical attenuation coefficient and the target attenuation coefficient. That is, correlation can reflect the degree of date correlation between the target date and historical dates. Specifically, correlation is a numerical value; for example, the value range of correlation can be [-1, 1]. The value of correlation can reflect the degree of date correlation; the larger the value, the greater the degree of date correlation. Correlation comparison refers to comparing the correlation similarity between the historical attenuation coefficient and the target attenuation coefficient. Specifically, cosine similarity can be used, or other methods such as Euclidean distance and modified cosine similarity can be used for correlation comparison.

[0122] This application's embodiments, by calculating the difference between the daily attenuation sub-parameter and historical date offset value, and the weekly attenuation sub-parameter and historical weekly offset value, take into account the time attenuation factor, thus minimizing the impact of historical dates distant from the target date on the prediction results. Secondly, by calculating the attenuation function based on the date attenuation function for the historical daily and weekly attenuation differences, a historical attenuation coefficient corresponding to each historical date can be obtained, thereby quantifying the influence of historical dates. Finally, by introducing the historical attenuation coefficient, it can be ensured that historical dates closer to the target date have a greater impact on the prediction results, thereby improving the accuracy of the parcel processing volume prediction model.

[0123] In step S103 of some embodiments, the first date type data refers to the type nature data of the target date. The second date type data refers to the type nature data of historical dates. For example, both the first and second date type data can be used to represent the day of the week, and can also be used to represent the day of a holiday. Date type similarity refers to the type nature similarity between the target date and historical dates. Date type similarity can be calculated using cosine similarity, or other methods such as Euclidean distance or modified cosine similarity.

[0124] Please see Figure 4 In some embodiments, step S103 may include, but is not limited to, steps S401 to S404:

[0125] Step S401: Extract the date type from the target date to obtain the target date type;

[0126] Step S402: Extract the date type for each historical date to obtain the historical date type for each historical date;

[0127] Step S403: Encode the target date type based on the preset date type encoding model to obtain the first date type data;

[0128] Step S404: Based on the date type encoding model, the historical date type of each historical date is encoded to obtain the second date type data of each historical date.

[0129] In step S401 of some embodiments, the target date type is used to characterize the type nature of the target date. For example, the target date type can be a weekday or a rest day. Date type extraction refers to the process of obtaining the target date type from the target date. Specifically, the target date type can be obtained by inputting the target date based on a pre-built date data model.

[0130] In step S402 of some embodiments, the historical date type is used to characterize the type nature of the historical date. The implementation method for date type extraction is the same as that of step S401 described above, and will not be repeated here. After obtaining the historical date type corresponding to a specific historical date, the historical date is updated to another historical date within a preset time range to obtain the historical date type corresponding to each historical date.

[0131] In step S403 of some embodiments, the date type encoding model is a pre-defined encoding model that can encode date types and obtain encoded data representing the date type. Date type encoding refers to obtaining date type data that uniquely corresponds to the target date type using a pre-defined date type encoding model. Specifically, the corresponding date type data can be obtained by inputting the date type into the date type encoding model. For example, the date type encoding model can be an encoding model constructed based on multi-hot combination, or it can be an encoding model constructed based on onehot. However, the onehot encoding model tends to isolate each date type, making its generalization ability for similar date types weak. For example, suppose we now distinguish different date types based on whether it is a weekday or not. Monday and Wednesday are both weekdays, while Saturday is a rest day. Then the date types "Monday" and "Wednesday" are more similar, while "Monday" and "Saturday" are less similar. However, it is difficult to distinguish between "Monday," "Wednesday," and "Saturday" using the one-hot encoding model, because the dot product of the vectors of "Monday" and "Wednesday" and "Monday" and "Saturday" is 0. Therefore, it is impossible to know that "Monday" and "Wednesday" are more similar than "Monday" and "Saturday." Therefore, this application can use the multi-hot combined encoding model as an example to further illustrate this application.

[0132] In a specific example, assuming lag0 is the target date type, and simultaneously taking the date types of the two days before and after it, the combined encoding yields a 5-dimensional vector, resulting in the first date type data (encoding rule: assuming 1 corresponds to a weekday / 2 corresponds to a 2-day holiday / 3 corresponds to a 3-day holiday / 4 corresponds to a long holiday...). The multi-lti-hot combined encoding model can be represented as follows:

[0133]

[0134] In the above model illustration, "Saturday" belongs to the "2-day holiday" type in the above coding rules, and is the first day of the "2-day holiday". If the target date type is "Saturday", "2-day holiday" corresponds to 2, then the code corresponding to lag0 is 2. Because lag0 and lag1 together can express "2-day holiday" and "Saturday" is the first day of the "2-day holiday", the code corresponding to lag1 is also 2. Moreover, the two days before "Saturday" are two consecutive working days, "working day" corresponds to 1, so the code corresponding to lag(-2) is 1, and the code corresponding to lag(-1) is 1. Finally, the day after "Saturday" is "Monday", "Monday" corresponds to a working day, so the code corresponding to lag2 is 1. Therefore, when the target date type of the target date is "Saturday", the corresponding first date type data can be expressed as [1, 1, 2, 2, 1]. Similarly, when the target date type of the target date is "Sunday", the corresponding first date type data can be expressed as [1, 2, 2, 1, 1]. When the target date type is "Wednesday", the corresponding first date type data can be expressed as [1, 1, 1, 1, 1].

[0135] It should be noted that the same logic applies to other target date types, based on the specific examples above. For instance, "Three-day holiday, Day 2" belongs to the "Three-day holiday" type in the above encoding rules. If the target date type is "Three-day holiday, Day 2", "Three-day holiday" corresponds to 3, so the encoding of lag0 is 3. Because lag(-1), lag0, and lag1 together express "Three-day holiday" and "Three-day holiday, Day 2" is the second day of "Three-day holiday", the encoding of lag(-1) is 3, and the encoding of lag1 is also 3. Moreover, the two days before "Three-day holiday, Day 2" are working days, and "working day" corresponds to 1, so the encoding of lag(-2) is 1. Finally, the second day after "Three-day holiday, Day 2" is a working day, so the encoding of lag2 is 1. Therefore, when the target date type is "Three-day holiday, Day 2", the corresponding first date type data can be expressed as [1, 3, 3, 3, 1]. Similarly, when the target date type is "last day of the long holiday", the corresponding first date type data can be expressed as [4, 4, 4, 1, 1].

[0136] In step S404 of some embodiments, the implementation of date type encoding is the same as that of step S403 described above, and will not be repeated here. After obtaining the second date type data corresponding to a specific historical date type, the historical date type is updated to the historical date type corresponding to another historical date within a preset time range, thereby obtaining the second date type data corresponding to each historical date.

[0137] This application embodiment encodes the target date type and the historical date type of each historical date using a preset date type encoding model. This converts the date types into a unified data representation, facilitating subsequent date type similarity calculations. Based on the first and second date type data, the date type similarity between each historical date and the target date is calculated. This quantifies the differences between the historical and target dates in terms of date type. By considering date type similarity, it ensures that the selected historical dates have a high degree of similarity to the target date in terms of date type, thereby improving the accuracy of the parcel processing volume prediction model.

[0138] In step S104 of some embodiments, the first package processing volume trend data refers to the pre-set package processing volume trend data for the target date. Specifically, the first package processing volume trend data can be the package processing volume of the previous stage. For example, in a logistics system, the package processing volume delivered to the destination node is unknown data. However, assuming that the package processing volume of the node preceding the target node is known, the package processing volume of the node preceding the target node can be used as the first package processing volume trend data. The second package processing volume trend data refers to the package processing volume trend data for historical dates. For a historical date, a second package processing volume trend data can be uniquely determined based on the historical package processing volume data. Specifically, the historical package processing volume data corresponding to this historical date can be used as the second package processing volume trend data for this historical date. According to this correspondence, the second package processing volume trend data corresponding to each historical date can be equal to the historical package processing volume data corresponding to each historical date. Trend similarity refers to the similarity between the target data and the historical date in terms of package processing volume trends. After obtaining the trend similarity for a specific historical date, the trend data of the second parcel processing volume for that specific historical date is updated with the trend data of the second parcel processing volume for another historical date within a preset time range, thus obtaining the trend similarity for each historical date. The trend similarity can be calculated using cosine similarity, or other methods such as Euclidean distance or modified cosine similarity.

[0139] Please see Figure 5 In some embodiments, step S104 may also include, but is not limited to, steps S501 to S503:

[0140] Step S501: Obtain the preset trend prediction model;

[0141] Step S502: Based on the trend prediction model, perform trend prediction on the first parcel processing volume trend data and the target date to obtain the target trend data;

[0142] Step S503: Based on the trend prediction model, perform trend prediction on the second parcel processing volume trend data and historical dates to obtain historical trend data.

[0143] In step S501 of some embodiments, the trend prediction model refers to a model that predicts the trend of target data and historical dates.

[0144] In step S502 of some embodiments, the target trend data refers to the trend data for the target date. In some embodiments, the trend prediction model includes a potential energy prediction sub-model and an energy level prediction sub-model. Specifically, step S502 includes: performing potential energy prediction on the first package processing volume trend data and the target date based on the potential energy prediction sub-model to obtain target potential energy data; performing energy level prediction on the first package processing volume trend data and the target date based on the energy level prediction sub-model to obtain target energy level data; and constructing target trend data based on the target potential energy data and the target energy level data.

[0145] It should be noted that the potential energy prediction sub-model refers to a model that predicts the potential energy based on the trend data of the target date and the first package processing volume. Target potential energy data refers to the potential energy data for the target date, which can reflect the fluctuation trend of the target date based on the trend data of the first package processing volume. Target potential energy data can reflect the position of the trend data of the first package processing volume for the target date within a preset period. For example, when the target potential energy data for the target date is greater than 1, the trend data of the first package processing volume for the target date is at a high level within a preset period; when the target potential energy data for the target date is less than 1, the trend data of the first package processing volume for the target date is at a low level within a preset period. The energy level prediction sub-model refers to a model that predicts the energy level based on the trend data of the target date and the first package processing volume. Target energy level data refers to the energy level data for the target date, which can reflect the fluctuation trend of the target date based on the trend data of the first package processing volume. Target energy level data can reflect the direction of trend change of the trend data of the first package processing volume for the target date. Furthermore, the target energy level data can be a vector. For example, when the target energy level data for the target date is greater than 0, the trend of the first package processing volume trend data for the target date is upward; when the target energy level data for the target date is less than 0, the trend of the first package processing volume trend data for the target date is downward. The target trend data is a two-dimensional data combination of target potential energy data and target energy level data. Specifically, the target potential energy data and target energy level data can be multiplied by preset weights respectively, and then combined to obtain the target trend data. The preset weights can be set according to the importance of the target potential energy data and target energy level data, and are not specifically limited here. For example, the preset weights can be 1:1, 1:2, etc. Therefore, the embodiments of this application can change the magnitude of the target potential energy data and target energy level data according to different preset weights to construct the target trend data.

[0146] It should be noted that the potential energy prediction sub-model can be a model constructed based on the following formula (1).

[0147]

[0148] Among them, X i Let i refer to the i-th day, assuming the i-th day is the target date, Q(X) i H(X) refers to the trend data of the first package processed on day i. i H(X) refers to the target potential energy data for day i. 2k (days) can represent a preset period, where k is a positive integer. This preset period is a small time interval within a preset time range; therefore, the preset period is smaller than the preset time range. For example, the preset time range could be 180 (days), and 2k could be 6 (days). In other words, the potential energy prediction sub-model can reflect the proportion of the trend data of the first package processing volume corresponding to the target date to the average of the trend data of the first package processing volume corresponding to the preset period of 2k days. If H(X) i If H(X) > 1, it means that the trend data of the first package processing volume corresponding to the target date is greater than the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days; if H(X) > 1, it means that the trend data of the first package processing volume corresponding to the target date is greater than the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days; i If H(X) = 1, it means that the trend data of the first package processing volume corresponding to the target date is equal to the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days; if H(X) = 1, it means that the trend data of the first package processing volume corresponding to the target date is equal to the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days; i If ) < 1, it means that the trend data of the first package processing volume corresponding to the target date is less than the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days.

[0149] It should be noted that the energy level prediction sub-model can be a model constructed based on the following formula (2):

[0150]

[0151] Wherein, D(X) i ) refers to the target energy level data for day i. In other words, the energy level prediction sub-model can reflect the difference between the trend data of the first parcel processing volume corresponding to the target date and the average value of the trend data of the first parcel processing volume corresponding to the preset period of 2k days. If D(X) i If D(X) > 0, it means that the trend data of the first parcel processing volume corresponding to the target date is greater than the average value of the trend data of the first parcel processing volume corresponding to the preset period of 2k days, and the trend direction is upward; if D(X) > 0, it means that the trend data of the first parcel processing volume corresponding to the target date is greater than the average value of the trend data of the first parcel processing volume corresponding to the preset period of 2k days, and the trend direction is upward. i If D(X) = 0, it means that the trend data of the first package processing volume corresponding to the target date is equal to the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days, and the trend change direction is unchanged; if D(X) = 0, it means that the trend data of the first package processing volume corresponding to the target date is equal to the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days, and the trend change direction is unchanged. iIf ) < 0, it means that the trend data of the first package processing volume corresponding to the target date is less than the average value of the trend data of the first package processing volume corresponding to the preset period of 2k days, and the trend change direction is downward.

[0152] In step S503 of some embodiments, historical trend data refers to trend data for historical dates. In some embodiments, historical trend data is obtained by predicting the trend data of the second parcel processing volume and the historical dates based on a trend prediction model, including: predicting the potential energy of the second parcel processing volume trend data and the historical dates based on a potential energy prediction sub-model to obtain historical potential energy data; predicting the energy level of the second parcel processing volume trend data and the historical dates based on an energy level prediction sub-model to obtain historical energy level data; and constructing historical trend data based on the historical potential energy data and the historical energy level data.

[0153] The specific method for solving historical potential energy data is the same as that for solving the target potential energy data in step S502 above, and will not be repeated here. The specific method for solving historical energy level data is the same as that for solving the target energy level data in step S502 above, and will not be repeated here. The specific method for constructing historical trend data is the same as that for constructing the target trend data in step S502 above, and will not be repeated here.

[0154] This application embodiment obtains more accurate trend information by performing trend prediction on the first parcel processing volume trend data and the target date, and the second parcel processing volume trend data and historical dates, thereby improving the accuracy of subsequent trend similarity calculations. By using a preset trend prediction model, the parcel processing volume prediction model can better adapt to different trend changes, thus improving the model's generalization ability. By calculating trend similarity, historical dates with high trend similarity to the target date can be selected from multiple historical dates, and these have significant reference value.

[0155] In step S105 of some embodiments, the target historical date refers to a historical date that meets the requirements in terms of correlation, date type similarity, and trend similarity. The parcel processing volume prediction model is a model that predicts the parcel processing volume for the target date.

[0156] Please see Figure 6 In some embodiments, step S105 includes, but is not limited to, steps S601 to S605:

[0157] Step S601: Obtain the pre-set first weight, second weight, and third weight; the first weight corresponds to the correlation between each historical date and the target date, the second weight corresponds to the date type similarity between each historical date and the target date, and the third weight corresponds to the trend similarity between each historical date and the target date.

[0158] Step S602: Calculate the correlation based on the first weight to obtain the first weighted data;

[0159] Step S603: Calculate the date type similarity based on the second weight to obtain the second weighted data;

[0160] Step S604: Calculate the trend similarity based on the third weight to obtain the third weighted data;

[0161] Step S605: Select dates within a preset time range based on the first weighted data, the second weighted data, and the third weighted data to determine multiple target historical dates.

[0162] In step S601 of some embodiments, the first weight refers to the relative importance of relevance. The second weight refers to the relative importance of date type similarity. The third weight refers to the relative importance of trend similarity. The relative order of the first, second, and third weights can be flexibly determined based on the actual level of importance. For example, if the importance of relevance, date type similarity, and trend similarity gradually increases when selecting target historical dates, then the corresponding third weight can be greater than the second weight, and the second weight can be greater than the first weight.

[0163] In step S602 of some embodiments, the first weighted data refers to the data obtained by multiplying the correlation by the first weight.

[0164] In step S603 of some embodiments, the second weighted data refers to the data obtained by multiplying the date type similarity by the second weight.

[0165] In step S604 of some embodiments, the third weighted data refers to the data obtained by multiplying the trend similarity by the third weight.

[0166] In step S605 of some embodiments, the target historical date refers to a historical date obtained based on the first weighted data, the second weighted data, and the third weighted data. Date selection refers to the process of selecting a target historical date based on the first weighted data, the second weighted data, the third weighted data, and a preset time range.

[0167] This application embodiment, by pre-setting a first weight, a second weight, and a third weight, allows for the allocation of different weights to relevance, date type similarity, and trend similarity according to actual needs, thereby achieving personalized relevance assessment. Through weighted calculation, relevance, date type similarity, and trend similarity can be combined to obtain a more comprehensive and accurate relevance assessment result, thus improving the screening quality of target historical dates. This method allows for adjustment of weights based on actual conditions, enabling the parcel processing volume prediction model to adapt to different application scenarios and data characteristics, improving the flexibility and adaptability of the parcel processing volume prediction model.

[0168] Please see Figure 7 In some embodiments, step S605 may include, but is not limited to, steps S701 to S703:

[0169] Step S701: For each historical date, sum the first weighted data, the second weighted data, and the third weighted data to obtain the candidate similarity for each historical date;

[0170] Step S702: Sort the candidate similarity of multiple historical dates within a preset time range in descending order to obtain the target historical date sequence;

[0171] Step S703: Select a target historical date from the target historical date sequence, which is a preset date threshold number.

[0172] In step S701 of some embodiments, the candidate similarity is used to characterize the similarity between the historical date and the target date. Data summation refers to the process of superimposing the first weighted data, the second weighted data, and the third weighted data. For a specific historical date, the first weighted data, the second weighted data, and the third weighted data are added together to obtain the candidate similarity for that specific historical date. After obtaining the candidate similarity for a specific historical date, the first weighted data, the second weighted data, and the third weighted data corresponding to this specific historical date are updated to the first weighted data, the second weighted data, and the third weighted data corresponding to another historical date within a preset time range, thus obtaining the candidate similarity for each historical date.

[0173] In step S702 of some embodiments, descending order means sorting from largest to smallest. The target historical date sequence refers to the sequence obtained by sorting the candidate similarities of multiple historical dates within a preset time range from largest to smallest. For example, if the preset time range can be 180 days, then arranging the candidate similarities corresponding to each historical date within 180 days from largest to smallest will result in a target historical date sequence that reflects the degree of similarity to the target date.

[0174] In step S703 of some embodiments, the target historical date refers to a historical date with a high similarity to the target date. The preset date threshold refers to a pre-set time threshold, and the number of preset date thresholds is used to characterize the number of target historical dates required. For example, if the number of preset date thresholds is 30, then 30 target historical dates are selected from the target historical date sequence in chronological order.

[0175] This embodiment of the application, by calculating the sum of first weighted data, second weighted data, and third weighted data, comprehensively considers multiple factors such as correlation, date type similarity, and trend similarity, thereby more comprehensively evaluating the candidate similarity between each historical date and the target date within a preset time range. Secondly, by sorting the candidate similarities in descending order, multiple target historical dates with high similarity to the target date can be selected from multiple historical dates within the preset time range. These target historical dates can provide more accurate and valuable information, helping to improve the accuracy of the parcel handling volume prediction model. Thirdly, by setting a preset date threshold, the number of target historical dates can be limited, thereby reducing the amount of data required for training the parcel handling volume prediction model and lowering computational complexity. Fourthly, by selecting multiple target historical dates with high similarity to the target date, the training time of the parcel handling volume prediction model can be shortened, improving the training efficiency. Finally, by training on historical parcel handling volume data corresponding to multiple target historical dates, the parcel handling volume prediction model can better adapt to different date types and trend changes, thereby improving the generalization ability of the parcel handling volume prediction model.

[0176] In one embodiment, preset fourth, fifth, and sixth weights are obtained. The fourth weight may correspond to the target attenuation coefficient. The fifth weight may correspond to the first date type data. The sixth weight may correspond to the target trend data. In this embodiment, the target attenuation coefficient is multiplied by the fourth weight, the first date type data is multiplied by the fifth weight, and the target trend data is multiplied by the sixth weight, and then a target multidimensional vector W is formed according to a preset combination. a For example, the target decay coefficient can be one-dimensional data, the first date type data can be five-dimensional data, and the target trend data can be two-dimensional data; therefore, the dimension of the target multidimensional vector can be 8. The fourth weight can correspond to the historical decay coefficient. The fifth weight can correspond to the second date type data. The sixth weight can correspond to the historical trend data. Similarly, the historical multidimensional vector W can be obtained. bSimilarly, the dimension of the historical multidimensional vector can be 8. After obtaining the target multidimensional vector and the historical multidimensional vector, the candidate similarity between the historical date and the target date is calculated based on the target multidimensional vector and the historical multidimensional vector. For example, if the cosine similarity method is used to calculate the candidate similarity between the historical date and the target date, the calculation process of the candidate similarity can be shown in the following formula (3):

[0177]

[0178] The above steps yield candidate similarities between a specific historical date and a target date. After obtaining these candidate similarities, the candidate similarities for each historical date within a preset time range are calculated. These candidate similarities are then arranged in descending order to obtain a candidate similarity list. A preset threshold number of target historical dates are selected in chronological order. This preset threshold number of target historical dates is then used as sample data for training the package processing volume prediction model.

[0179] After step S105 in some embodiments, after training the package processing volume prediction model, in practical applications, a preset date for the package processing volume to be predicted can be obtained; the preset date is then input into the package processing volume prediction model to obtain the package processing volume to be predicted. The preset date refers to a future date on which the package processing volume prediction needs to be performed.

[0180] The embodiments of this application pre-train a package processing volume prediction model using historical package processing volume data within a target date and preset time range. Based on the preset date and the package processing volume prediction model, the package processing volume to be predicted can be obtained directly.

[0181] Please see Figure 8 This application also provides a data processing apparatus that can implement the above-described data processing method. The apparatus includes:

[0182] The data acquisition module is used to acquire the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range;

[0183] The correlation calculation module is used to calculate the correlation between each historical date and the target date based on the nearest neighbor relationship between the target date and each historical date within a preset time range;

[0184] The type calculation module is used to obtain the first date type data of the target date and the second date type data of each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data;

[0185] The trend prediction module is used to obtain the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data.

[0186] The date determination module is used to determine multiple target historical dates within a preset time range based on correlation, date type similarity, and trend similarity, and to train a parcel processing volume prediction model based on the historical parcel processing volume data corresponding to the multiple target historical dates.

[0187] The specific implementation of this data processing device is basically the same as the specific embodiment of the data processing method described above, and will not be repeated here.

[0188] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described data processing method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0189] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0190] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0191] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the data processing method of the embodiments of this application.

[0192] The input / output interface 903 is used to implement information input and output;

[0193] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0194] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0195] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0196] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described data processing method.

[0197] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0198] This application provides a data processing method, apparatus, electronic device, and storage medium. Based on the proximity relationship, date type similarity, and trend similarity between the target date and each historical date, it selects multiple candidate historical dates with high similarity to the target date. Then, it trains a pre-trained model (such as a parcel volume prediction model) corresponding to the system of the resource application scenario using historical parcel processing volume data from the selected multiple target historical dates, thereby improving the recognition accuracy of the pre-trained model. In other words, this application addresses the bias in existing solutions (directly selecting based on the closest data partitions of the same date type) by expressing the data partition similarity problem as a similarity problem between low-dimensional vectors under three constraints. The data processing method of this application has strong interpretability, as the timeliness effect is recent, the date types are similar, and the future trends are similar. This application also improves the efficiency of the pre-trained model. For example, this application can compress the annual sample size of 36 billion to about 100 million (the upper limit of the current training data for the parcel volume prediction model), effectively reducing the model training cost while ensuring the performance of the parcel volume prediction model.

[0199] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0200] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0202] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0203] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0204] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0205] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0206] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0207] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0208] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0209] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range; Based on the nearest neighbor relationship between the target date and each historical date within the preset time range, the correlation between each historical date and the target date is calculated; Obtain first date type data for the target date and second date type data for each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data; Obtain the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data; Based on the correlation, the date type similarity, and the trend similarity, multiple target historical dates are determined within the preset time range, and a parcel processing volume prediction model is trained based on the historical parcel processing volume data corresponding to the multiple target historical dates.

2. The method according to claim 1, characterized in that, The determination of multiple target historical dates within the preset time range based on the correlation, date type similarity, and trend similarity includes: Obtain a pre-set first weight, second weight, and third weight; the first weight corresponds to the correlation between each historical date and the target date, the second weight corresponds to the date type similarity between each historical date and the target date, and the third weight corresponds to the trend similarity between each historical date and the target date; The correlation is weighted based on the first weight to obtain the first weighted data; The date type similarity is weighted based on the second weight to obtain the second weighted data; The trend similarity is weighted based on the third weight to obtain the third weighted data. Based on the first weighted data, the second weighted data, and the third weighted data, a date selection is performed within the preset time range to determine multiple target historical dates.

3. The method according to claim 2, characterized in that, The step of selecting dates within the preset time range based on the first weighted data, the second weighted data, and the third weighted data to determine multiple target historical dates includes: For each historical date, the first weighted data, the second weighted data, and the third weighted data are summed to obtain the candidate similarity for each historical date; The candidate similarities of multiple historical dates within the preset time range are sorted in descending order to obtain the target historical date sequence; Select a preset date threshold number of target historical dates from the target historical date sequence.

4. The method according to claim 1, characterized in that, The step of calculating the correlation between each historical date and the target date based on the nearest neighbor relationship between the target date and each historical date within the preset time range includes: Obtain the preset offset reference data and date decay parameters; The historical date offset value is determined based on the offset reference data and the historical date; Determine the target date offset value based on the historical date and the target date; Based on a preset date decay function and the date decay parameters, the historical date offset value is decayed to obtain a historical decay coefficient; Based on the date decay function and the date decay parameter, the target date offset value is decayed to obtain the target decay coefficient; The correlation is determined by comparing the historical attenuation coefficient and the target attenuation coefficient.

5. The method according to claim 4, characterized in that, The date decay parameter includes a daily decay sub-parameter and a weekly decay sub-parameter. The process of performing date decay on the historical date offset value based on a preset date decay function and the date decay parameter to obtain a historical decay coefficient includes: The difference between the day decay sub-parameter and the historical date offset value is calculated to obtain the historical day decay difference. The historical week offset value is determined based on the historical date offset value; The difference between the cycle attenuation sub-parameter and the historical cycle offset value is calculated to obtain the historical cycle attenuation difference. The historical attenuation coefficient is obtained by calculating the attenuation function based on the historical daily attenuation difference and the historical weekly attenuation difference.

6. The method according to claim 1, characterized in that, The process of obtaining the first date type data of the target date and the second date type data of each of the historical dates includes: Extract the date type from the target date to obtain the target date type; For each historical date, extract the date type to obtain the historical date type for each historical date; The target date type is encoded based on a preset date type encoding model to obtain the first date type data; Based on the date type encoding model, the historical date type of each historical date is encoded to obtain the second date type data of each historical date.

7. The method according to claim 1, characterized in that, The step of obtaining the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each of the historical dates includes: Obtain the preset trend prediction model; Based on the trend prediction model, the trend data of the first package processing volume and the target date are predicted to obtain the target trend data; Based on the trend prediction model, the trend data of the second parcel processing volume and the historical date are used to predict the trend, and historical trend data is obtained.

8. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire the target date for the predicted parcel processing volume and historical parcel processing volume data within a preset time range; The correlation calculation module is used to calculate the correlation between each historical date and the target date based on the nearest neighbor relationship between the target date and each historical date within the preset time range; A type calculation module is used to obtain first date type data of the target date and second date type data of each historical date, and calculate the date type similarity between each historical date and the target date based on the first date type data and the second date type data; The trend prediction module is used to acquire the first parcel processing volume trend data corresponding to the target date and the second parcel processing volume trend data corresponding to each historical date, and calculate the trend similarity between each historical date and the target date based on the first parcel processing volume trend data and the second parcel processing volume trend data. The date determination module is used to determine multiple target historical dates within the preset time range based on the correlation, the date type similarity, and the trend similarity, and to train a parcel processing volume prediction model based on the historical parcel processing volume data corresponding to the multiple target historical dates.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a data processing method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Strategy information generation method, system and device, storage medium and program product

    CN121504224A