Logistics data processing method and device, storage medium and electronic equipment

By obtaining the shipping frequency of the logistics path, dividing the initial data set and calculating the data set equalization coefficient, adaptively diversion of the residual logistics path, the problem of random composition of logistics data is solved, and the data distribution equalization and testing effect are improved.

CN120373997APending Publication Date: 2025-07-25BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410108730.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the data groups are randomly composed after the logistics data are grouped, resulting in poor logistics strategy testing results.

Method used

By obtaining the shipping frequency of the logistics path, determining the target logistics path and dividing it into the initial data set, calculating the data set equalization coefficient, adaptively equalizing diversion of the remaining logistics path to the initial data set, forming a target data set with balanced data distribution.

Benefits of technology

The target data set with balanced data distribution is achieved, and the effectiveness and accuracy of logistics strategy testing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373997A_ABST
    Figure CN120373997A_ABST
Patent Text Reader

Abstract

The invention provides a logistics data processing method and device, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the following steps: acquiring a plurality of logistics paths and the delivery frequency of each logistics path; determining a target logistics path from the plurality of logistics paths according to the delivery frequency of each logistics path, and dividing the target logistics path into a first number of initial data sets; for each remaining logistics path except the target logistics path, sequentially calculating a data set equalization coefficient after the remaining logistics paths are added into each initial data set; respectively adding each residual logistics path into one of the initial data sets according to the data set equalization coefficient so as to obtain a first number of target data sets; according to the method, self-adaptive balanced distribution can be automatically carried out on each residual logistics path based on the data set balance coefficient, and a target data set with balanced data distribution is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the logistics scenario, there are various logistics strategies applied to the production environment. These logistics strategies often need to be improved. Therefore, it is necessary to group the existing logistics data so as to use the logistics strategies or improved logistics strategies to test the grouped logistics data and obtain the effects of various logistics strategies.

[0003] In the related art, the existing logistics data is usually grouped by using a random shunt algorithm based on hash shunting. Therefore, the composition of the data in each group after grouping is very random, often resulting in a large impact on the test effect of the logistics strategy and a poor test effect.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a logistics data processing method, device, electronic device and storage medium, which can automatically perform adaptive balanced shunting on each remaining logistics path based on the dataset balance coefficient to obtain a target dataset with balanced data distribution.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0007] According to one aspect of the present disclosure, a logistics data processing method is provided, including: obtaining a plurality of logistics paths and the shipping frequencies of each logistics path; determining a target logistics path from the plurality of logistics paths according to the shipping frequencies of each logistics path, and dividing the target logistics path into a first number of initial datasets; for each remaining logistics path except the target logistics path, calculating the dataset balance coefficient after adding the remaining logistics path to each initial dataset in turn; adding each remaining logistics path to one of the initial datasets according to the dataset balance coefficient to obtain a first number of target datasets.

[0008] In one embodiment of the present disclosure, the dataset balance coefficient includes a sample size balance coefficient; wherein, for each remaining logistics path other than the target logistics path, the dataset balance coefficient after adding the remaining logistics path to each initial dataset is calculated in sequence, including: determining the target remaining logistics path as the remaining logistics path with the highest shipping frequency among the remaining logistics paths; determining the current sample size balance coefficient according to the logistics paths in all initial datasets; if the current sample size balance coefficient is greater than the balance coefficient threshold, calculating the sample size balance coefficient after adding the target remaining logistics path to each initial dataset; wherein, adding each remaining logistics path to one of the initial datasets according to the dataset balance coefficient includes: adding the target remaining logistics path to the initial dataset corresponding to the smallest sample size balance coefficient.

[0009] In one embodiment of the present disclosure, the dataset balance coefficient further includes a feature distribution difference coefficient; wherein, for each remaining logistics path other than the target logistics path, the dataset balance coefficient after adding the remaining logistics path to each initial dataset is calculated in sequence, further including: if the current sample size balance coefficient is less than or equal to the balance coefficient threshold, calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial dataset; wherein, adding each remaining logistics path to one of the initial datasets according to the dataset balance coefficient further includes: adding the target remaining logistics path to the initial dataset corresponding to the smallest feature distribution difference coefficient.

[0010] In one embodiment of the present disclosure, determining the current sample size balance coefficient according to the logistics paths in all initial datasets includes: determining the shipping days of the logistics paths in each initial dataset within the historical period according to the shipping frequencies of the logistics paths in each initial dataset; determining the sample size included in each initial dataset according to the shipping days of the logistics paths in each initial dataset; determining the current sample size balance coefficient according to the sample sizes included in all initial datasets.

[0011] In one embodiment of the present disclosure, calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial data set includes: for each target initial data set, adding the target remaining logistics path to the target initial data set to obtain a target intermediate data set; determining the attribute feature distribution information of the target intermediate data set and each other initial data set according to the attribute values of the logistics paths in the target intermediate data set and each other initial data set under a preset attribute; the other initial data sets are the initial data sets other than the target initial data set; determining the feature distribution distance between the target intermediate data set and each other initial data set according to the attribute feature distribution information of the target intermediate data set and each other initial data set; and performing a weighted average on the feature distribution distances between the target intermediate data set and each other initial data set to obtain the feature distribution difference coefficient of the target intermediate data set.

[0012] In one embodiment of the present disclosure, determining a target logistics path from multiple logistics paths according to the shipping frequency of each logistics path and dividing the target logistics path into a first number of initial data sets includes: sorting the multiple logistics paths in descending order according to the shipping frequency of each logistics path, and determining the first second number of logistics paths as the target logistics path; wherein, the second number is an integer multiple of the first number; and evenly and randomly dividing the target logistics path into a first number of initial data sets; wherein, the number of target logistics paths included in each initial data set is the same.

[0013] In one embodiment of the present disclosure, the logistics data processing method further includes: obtaining a first number of logistics strategies to be tested; matching the first number of target data sets with the first number of logistics strategies to be tested one by one to obtain a first number of test combinations; applying the logistics strategy to be tested in each test combination to the target data set to obtain the test effect corresponding to the test combination; and using the logistics strategy to be tested in the test combination with the best test effect as the target logistics strategy to apply the target logistics strategy to each logistics path.

[0014] According to another aspect of the present disclosure, a logistics data processing device is provided, including: an acquisition module, configured to acquire multiple logistics paths and the shipping frequency of each logistics path; an initial data set division module, configured to determine a target logistics path from the multiple logistics paths according to the shipping frequency of each logistics path and divide the target logistics path into a first number of initial data sets; a calculation module, configured to calculate, for each remaining logistics path other than the target logistics path, the data set balance coefficient after adding the remaining logistics path to each initial data set in sequence; and a target data set determination module, configured to add each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient to obtain a first number of target data sets.

[0015] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the above-described logistics data processing method.

[0016] According to yet another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above-described logistics data processing method by executing the executable instructions.

[0017] On the one hand, the logistics data processing method provided by the embodiments of the present disclosure can implement an initialization method of an initial data set with the shipping frequency as a consideration factor, providing a cold start solution for the logistics data processing method provided by the present disclosure; on the other hand, it can automatically process each remaining logistics path based on the data set balance coefficient and adaptively balance and divert each remaining logistics path. Therefore, all logistics paths can be processed while ensuring the overall balance of the initial data set, obtaining a target data set with balanced data distribution, and then the target data set with balanced data distribution can be further applied to the logistics production environment to provide a data basis for various logistics production requirements (such as strategy testing requirements).

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0020] Figure 1 A schematic diagram showing an exemplary system architecture to which the logistics data processing method of the embodiments of the present disclosure can be applied;

[0021] Figure 2 A flowchart showing the logistics data processing method of an embodiment of the present disclosure;

[0022] Figure 3 A schematic diagram showing the system interaction for implementing the logistics data processing method of an embodiment of the present disclosure;

[0023] Figure 4 A schematic diagram showing the process of the logistics data processing method of an embodiment of the present disclosure;

[0024] Figure 5A schematic diagram showing the result of feature distribution obtained in the logistics data processing method according to an embodiment of the present disclosure;

[0025] Figure 6 A block diagram showing a logistics data processing apparatus according to an embodiment of the present disclosure; and

[0026] Figure 7 A structural block diagram showing a logistics data processing computer device according to an embodiment of the present disclosure. Detailed implementation manners

[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0028] In addition, the drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0029] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0030] Figure 1 A schematic diagram showing an exemplary system architecture to which the logistics data processing method according to an embodiment of the present disclosure can be applied.

[0031] As Figure 1 shown, the system architecture may include a server 101, a network 102, and a client 103. The network 102 is used to provide a medium for a communication link between the client 103 and the server 101. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0032] In an exemplary embodiment, the client 103 for data transmission with the server 101 may include, but is not limited to, electronic devices of types such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, AR (Augmented Reality) devices, VR (Virtual Reality) devices, smart wearable devices, etc. Optionally, the operating systems running on the electronic devices may include, but are not limited to, Android system, IOS system, Linux system, Windows system, etc.

[0033] The server 101 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some practical applications, the server 101 may also be the server of a network platform, and the network platform may be, for example, a trading platform, a live broadcast platform, a social platform, or a music platform, etc., and the embodiments of the present disclosure do not limit this. Among them, the server may be a single server or a cluster formed by multiple servers, and the present disclosure does not limit the specific architecture of the server.

[0034] In an exemplary embodiment, the user may initiate an instruction for indicating the processing of logistics data to the server 101 through the client 103. The instruction may carry a first quantity, or may carry the type or acquisition channel of the logistics data to be processed; after receiving the instruction, the server 101 may execute the logistics data processing method provided by the present disclosure according to the instruction, and then obtain a first quantity of target data sets.

[0035] In an exemplary embodiment, the server 101 may also return the first quantity of target data sets to the client 103 for application in the logistics production environment according to the first quantity of target data sets. For example, the user may continue to initiate an application instruction to the server 101 through the client 103, and the application instruction may indicate the acquisition method of the first quantity of logistics strategies to be tested; after receiving the application instruction, the server 101 may obtain the first quantity of logistics strategies to be tested, and then conduct an AB test (a control variable-based controlled experiment method) of the logistics strategies according to the first quantity of logistics strategies to be tested and the first quantity of target data sets, and then select the target logistics strategy finally applied to actual production according to the experimental results.

[0036] In an exemplary embodiment, the process by which the server 101 is used to implement the logistics data processing method may be as follows: The server 101 obtains a plurality of logistics paths and the shipping frequencies of each logistics path; The server 101 determines a target logistics path from the plurality of logistics paths according to the shipping frequencies of each logistics path, and divides the target logistics path into a first number of initial data sets; For each remaining logistics path except the target logistics path, the server 101 sequentially calculates the data set balance coefficient after adding the remaining logistics path to each initial data set; The server 101 adds each remaining logistics path to one of the initial data sets according to the data set balance coefficient to obtain a first number of target data sets.

[0037] In addition, it should be noted that Figure 1 What is shown is only an application environment of the logistics data processing method provided by the present disclosure. Figure 1 The numbers of the server 101, the network 102, and the client 103 in

[0038] are merely illustrative. According to actual needs, there can be any number of clients, networks, and servers.

[0039] Figure 2 shows a flowchart of the logistics data processing method according to an embodiment of the present disclosure. The method provided by the embodiments of the present disclosure can be executed by a server 101 or a client 103 as shown in Figure 1 but the present disclosure is not limited thereto.

[0040] In the following illustrative examples, the server 101 is used as the execution subject for illustrative purposes. As shown in Figure 2 the logistics data processing method provided by the embodiments of the present disclosure may include the following steps.

[0041] Step S201, obtain a plurality of logistics paths and the shipping frequencies of each logistics path.

[0042] In this step, the logistics path may be a pre-planned line for logistics transportation in a logistics scenario, such as a railway line, a motor vehicle line, a non-motor vehicle line, or a combined line of at least two of these lines. The shipping frequency may be the frequency at which the logistics path is arranged to transport equipment for goods, and can be set for each logistics path based on actual production requirements or production environment; for example, the shipping frequency may be shipping goods every day, shipping goods 5 days a week, shipping goods 10 days a month, shipping goods every other day, etc.

[0043] In some practical applications, a time period (such as 30 days, 60 days, 180 days, etc.) can also be set. Then, the shipping days of each logistics path within this time period are counted from the historical shipping data, and the shipping frequency of each logistics path is determined based on the shipping days and the time period. Among them, the shipping days can refer to the days when there are goods shipped on the corresponding logistics path.

[0044] Step S203: Determine the target logistics path from multiple logistics paths according to the shipping frequency of each logistics path, and divide the target logistics path into the first number of initial data sets.

[0045] In this step, some logistics paths with higher shipping frequencies can be selected first as the target logistics paths, and then the first number of initial data sets are formed according to the selected target logistics paths. This step can be regarded as a way to initialize the initial data set, which can provide a cold start method for the calculation step of "calculating the initial data set to be added for each remaining logistics path" in the subsequent solution, so that the initial data set used in the subsequent calculation steps has a data basis.

[0046] Among them, the first number can be preset, and the value of the first number is also the number of target data sets that can be finally obtained. In some practical applications, the first number can be set according to the actual strategy test requirements; for example, if 4 logistics strategies to be tested are used in the strategy test requirements, then the first number can be a value greater than or equal to 4, and then the logistics data processing method in the embodiments of the present disclosure can be used to finally obtain more than or equal to 4 target data sets to ensure that the 4 logistics strategies to be tested are respectively matched with different target data sets, so as to ensure the smooth execution of the AB test, and thus the target logistics strategy that can be applied to actual production can be selected for the strategy test.

[0047] In some embodiments, step S203 can further include: sorting the multiple logistics paths in descending order according to the shipping frequency of each logistics path, and determining the first second number of logistics paths as the target logistics paths; where the second number is an integer multiple of the first number; dividing the target logistics path evenly and randomly into the first number of initial data sets; where the number of target logistics paths included in each initial data set is the same.

[0048] In this embodiment, the logistics paths with higher shipping frequencies can be selected to form an initial data set. Among them, the logistics data generated on a logistics path within a specified time period can be used as sample data in the data set. For example, the logistics data of a logistics path in one day can be used as one sample. Obviously, the logistics paths with higher shipping frequencies will generate more logistics data and a larger number of relevant samples within the same time period. Based on this, generating an initial data set according to the logistics paths with higher shipping frequencies can achieve the effect of "obtaining a larger number of samples in the initial data set through a smaller number of logistics paths", so that an initial data set with a large number of samples can be obtained efficiently. In addition, since the number of target logistics paths included in each initial data set obtained in this embodiment is the same and the shipping frequencies of the logistics paths are all relatively high, it can also ensure that the number of samples in each initial data set is basically the same and relatively large. Among them, the closer the number of samples in each initial data set is, the smaller the impact on the distribution differences of the data characteristics (including the characteristic of the number of samples) of each data set will be. Therefore, it is beneficial to make the division of the data set more average and stable, and thus it is more likely to obtain multiple data sets with similar data characteristics. Among them, if the shipping frequencies of all target logistics paths are the same (such as shipping every day), then the number of samples in each initial data set can also be guaranteed to be the same. It can be seen that through this embodiment, a smaller number of logistics paths can be used to make the number of samples in the initial data set larger and the number of samples equivalent, which not only realizes efficient data initialization but also is beneficial to the stability of data set division.

[0049] In this embodiment, the integer multiple can be a preset value, such as 2, 3, or 5, etc. The integer multiple can determine the number of logistics paths in the initial data set obtained in this step. For example, assuming the first number is K and 3K target logistics paths are selected, then the number of target logistics paths included in each initial data set obtained in this step should be 3.

[0050] In some practical applications, the integer multiple can be determined based on the number of logistics paths with higher shipping frequencies (such as the shipping frequency being greater than or equal to the frequency threshold) and the first number. For example, the quotient obtained by dividing the number of logistics paths with higher shipping frequencies by the first number can be calculated first, and then a positive integer less than or equal to this quotient can be selected as the integer multiple to ensure that the selected target logistics paths have relatively high shipping frequencies and their number (i.e., the second number) is sufficient to be evenly divided into the first number of initial data sets.

[0051] In some practical applications, an empty initial dataset can be created first, and the path identification information of each target logistics path can be obtained. Then, the random hash splitting algorithm is used to process each path identification information to evenly and randomly divide each target logistics path into the corresponding initial datasets. Among them, the path identification information can be the ODT information of the logistics path. This ODT information is the unique key value representing the logistics path in the logistics main-branch scheduling scenario, and it consists of 4 parts: the starting station of the line, the arrival station of the line, the departure time of the line, and the line network type (such as Network B, Network C, Network TC). These four parts can be connected by an underscore '_' to form the ODT information of the logistics path.

[0052] Step S205: For each remaining logistics path except the target logistics path, calculate the dataset balance coefficient after adding the remaining logistics path to each initial dataset in turn.

[0053] In this step, the dataset balance coefficient can be used to quantify the balance of the overall data feature distribution among the initial datasets. The dataset balance coefficient can be calculated jointly based on the information of all initial datasets, or can be calculated based on the information of one initial dataset relative to other initial datasets. The present disclosure does not make a limitation on this.

[0054] In some practical applications, the smaller the dataset balance coefficient, the more balanced the overall data feature distribution among the current all initial datasets. For example, the dataset balance coefficient can be a value calculated based on the Wassertein Distance (a characteristic distribution distance that measures the data distribution between two datasets).

[0055] Step S207: Add each remaining logistics path to one of the initial datasets according to the dataset balance coefficient to obtain a first number of target datasets.

[0056] As described above, taking "the smaller the dataset balance coefficient, the more balanced the overall data feature distribution among the current all initial datasets" as an example, after calculating the dataset balance coefficient after adding each remaining logistics path to each initial dataset each time, the remaining logistics path can be added to the initial dataset corresponding to the smallest dataset balance coefficient. For example, if the dataset balance coefficient is calculated jointly based on the information of all initial datasets, the remaining logistics path can be added to the initial dataset that generates the smallest dataset balance coefficient after adding; if the dataset balance coefficient is calculated based on the information of one initial dataset relative to other initial datasets, the remaining logistics path can be added to the initial dataset with the smallest dataset balance coefficient after adding.

[0057] By processing each remaining logistics path in this way successively, all logistics paths can be processed while ensuring the overall balance of the initial data set, and the first number of target data sets with balanced data distribution and very close various features can be obtained. It can be seen that this step realizes the automatic adaptive balanced diversion of each remaining logistics path.

[0058] In some embodiments, the logistics data processing method may further include: obtaining the first number of logistics strategies to be tested; matching the first number of target data sets with the first number of logistics strategies to be tested one by one to obtain the first number of test combinations; applying the logistics strategies to be tested in each test combination to the target data set to obtain the test effects corresponding to the test combinations; and taking the logistics strategy to be tested in the test combination with the best test effect as the target logistics strategy to apply the target logistics strategy to each logistics path.

[0059] Among them, since the data distribution in each target data set is balanced and various features are very close, when applying the logistics strategies to be tested in each test combination to the target data set, it can be considered that the target data sets in different test combinations are the same, and there are only differences in the logistics strategies to be tested. Therefore, the effect of the control experiment can be ensured to be good, and further, it can be ensured that the logistics strategy to be tested in the test combination with the best test effect can have the best effect in actual application.

[0060] Through the logistics data processing method provided by the present disclosure, the target logistics paths can be selected first according to the shipping frequencies of multiple logistics paths, and then the selected target logistics paths can be divided into the first number of initial data sets. Then, the data set balance coefficient after adding each remaining logistics path to each initial data set can be calculated, and further, each remaining logistics path can be automatically and adaptively diverted based on the data set balance coefficient, and the balance between all initial data sets after diversion can be ensured, so as to obtain the first number of target data sets with balanced data distribution. It can be seen that on the one hand, this solution selects the target logistics paths according to the shipping frequencies of multiple logistics paths and generates the first number of initial data sets based on this, and can realize an initialization method of the initial data set with the shipping frequency as a consideration factor, providing a cold start solution for the logistics data processing method provided by the present disclosure; on the other hand, this solution can automatically process each remaining logistics path based on the data set balance coefficient, and adaptively balance and divert each remaining logistics path. Therefore, all logistics paths can be processed while ensuring the overall balance of the initial data set, and the target data sets with balanced data distribution can be obtained. Furthermore, the target data sets with balanced data distribution can be further applied to the logistics production environment to provide a data basis for various logistics production requirements (such as strategy test requirements).

[0061] Figure 3The figure shows a schematic diagram of system interaction for implementing a logistics data processing method according to an embodiment of the present disclosure, as Figure 3 shown, including a database management system 301 (such as Mysql), a data warehouse tool 302 (Hive), an adaptive shunting algorithm system 303, a data warehouse tool 302 (Hive), and a file system 304 (HDFS).

[0062] Among them, the database management system 301 (such as Mysql) and the data warehouse tool 302 (Hive) may store the original information of multiple logistics paths, such as the starting station, the destination station, the estimated departure time, the line network attributes (such as Network B, Network C, Network TC), etc.; the adaptive shunting algorithm system 303 may be a system capable of executing the logistics data processing method provided by the present disclosure; the data warehouse tool 302 (Hive) and the file system 304 (HDFS) may be used to store the results generated by the adaptive shunting algorithm system 303, for example, the first number of target data sets obtained in this method may be stored.

[0063] Refer to Figure 3 , in the system interaction process of implementing the logistics data processing method, it may include a data extraction stage, a line shunting stage, and a shunting result saving stage. Specifically, in the data extraction stage, the adaptive shunting algorithm system 303 may obtain the original information of multiple logistics paths from the database management system 301 (such as Mysql) and / or the data warehouse tool 302 (Hive), and specifically may obtain multiple logistics paths and the shipping frequencies of each logistics path. In the line shunting stage, the adaptive shunting algorithm system 303 may execute the logistics data processing method provided by the present disclosure to obtain the first number of target data sets. In the shunting result saving stage, the adaptive shunting algorithm system 303 may send the first number of target data sets to the data warehouse tool 302 (Hive) and / or the file system 304 (HDFS) so that the data warehouse tool 302 (Hive) and / or the file system 304 (HDFS) store the first number of target data sets for other logistics production requirements.

[0064] In some embodiments, the dataset balance coefficient includes a sample size balance coefficient.

[0065] The sample size balance coefficient may be a quantization coefficient reflecting the sample size balance of all initial data sets.

[0066] In some practical applications, the larger the sample size balance coefficient, the greater the difference in sample size among all initial data sets at this time; the smaller the sample size balance coefficient, the smaller the difference in sample size among all initial data sets at this time, that is, the more balanced. Based on this, step S205 may include: determining the target remaining logistics path as the remaining logistics path with the highest shipping frequency in the remaining logistics paths; determining the current sample size balance coefficient according to the logistics paths in all initial data sets; if the current sample size balance coefficient is greater than the balance coefficient threshold, calculating the sample size balance coefficient after adding the target remaining logistics path to each initial data set.

[0067] In this step, when each remaining logistics path needs to be processed, the remaining logistics path with the highest current shipping frequency can be processed first. Among them, when diverting a logistics path (that is, allocating it to a certain initial data set), in fact, the sample size related to this logistics path is uniformly diverted. As mentioned above, the larger the shipping frequency of the remaining logistics path, the more samples are related to it. Therefore, by giving priority to processing the remaining logistics path with a high shipping frequency, the large amount of sample size corresponding to this remaining logistics path can be arranged first, thus avoiding the problem of unbalanced sample size allocation caused by processing these large amounts of sample size late.

[0068] Among them, it is also possible to first sort all the remaining logistics paths according to the shipping frequency from high to low, and then process each remaining logistics path in the sorting in turn.

[0069] If the current sample size balance coefficient is greater than the balance coefficient threshold, it means that the initial data sets are not balanced enough in terms of sample size at this time. Therefore, the balance of the sample size can be adjusted first according to this embodiment.

[0070] In some embodiments, determining the current sample size balance coefficient according to the logistics paths in all initial data sets includes: determining the shipping days of the logistics paths in each initial data set within the historical period according to the shipping frequency of the logistics paths in each initial data set; determining the sample size included in each initial data set according to the shipping days of the logistics paths in each initial data set; determining the current sample size balance coefficient according to the sample size included in all initial data sets.

[0071] In some practical applications, the determination method of the sample size balance coefficient can be:

[0072]

[0073] Among them, K represents the first quantity, that is, the number of the initial quantity sets; p k represents the proportion of the sample size in the kth initial quantity set in the total sample size, where the total sample size is the sum of the sample sizes in all initial quantity sets. According to this determination method, if each p kThe greater the difference, the greater B(p); if each p k has a smaller difference, then B(p) is smaller.

[0074] In this embodiment, the balance coefficient threshold can be determined according to the first quantity (i.e., the quantity of the initial data set). For example, if the first quantity is 2, obviously the minimum value of B(p) is 0.5, then the balance coefficient threshold can be set to a value greater than 0.5 (such as 0.6, 0.7, etc.); if the first quantity is 4, obviously the minimum value of B(p) is 0.25, then the balance coefficient threshold can be set to a value greater than 0.25 (such as 0.3, 0.4, etc.).

[0075] Similarly, in some embodiments, calculating the sample volume balance coefficient after adding the target remaining logistics path to each initial data set may include:

[0076] For each target initial data set, add the target remaining logistics path to the target initial data set to obtain a target intermediate data set; according to the shipping frequencies of the logistics paths in the target intermediate data set and each other initial data set, determine the shipping days of the logistics paths in the target intermediate data set and each other initial data set within the historical period; where the other initial data sets are the initial data sets except the target initial data set; then respectively determine the sample volume included in the target intermediate data set and the sample volumes included in each other initial data set according to the shipping days of the logistics paths in the target intermediate data set and each other initial data set; finally, determine the sample volume balance coefficient after adding the target remaining logistics path to the target initial data set according to the sample volume included in the target intermediate data set and all other initial data sets.

[0077] In this way, the sample volume balance coefficients after adding the target remaining logistics path to each initial data set can be obtained.

[0078] Further, step S207 may include: adding the target remaining logistics path to the initial data set corresponding to the smallest sample volume balance coefficient.

[0079] That is, when the current sample volume balance coefficient is greater than the balance coefficient threshold, add the target remaining logistics path to the initial data set corresponding to the smallest sample volume balance coefficient.

[0080] The smaller the sample volume balance coefficient, the more it can illustrate that the sample volumes among the initial data sets after addition are more balanced. Therefore, adding the target remaining logistics path to the initial data set corresponding to the smallest sample volume balance coefficient can make the sample volumes among all initial data sets tend to be balanced.

[0081] In some embodiments, the data set balance coefficient further includes a feature distribution difference coefficient.

[0082] The feature distribution difference coefficient can be a quantization coefficient for the feature distribution difference between a certain initial data set and other initial data sets.

[0083] In some practical applications, the larger the feature distribution difference coefficient is, the greater the feature distribution difference between an initial data set and other initial data sets; the smaller the feature distribution difference coefficient is, the smaller the feature distribution difference between an initial data set and other initial data sets. Based on this, step S205 may further include: if the current sample size balance coefficient is less than or equal to the balance coefficient threshold, calculate the feature distribution difference coefficient after adding the target remaining logistics path to each initial data set.

[0084] When the current sample size balance coefficient is less than or equal to the balance coefficient threshold, it indicates that the initial data sets can be considered balanced in terms of the sample size dimension at this time. Therefore, the balance of the initial data sets in terms of features in other dimensions can be adjusted.

[0085] In some embodiments, calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial data set includes: for each target initial data set, adding the target remaining logistics path to the target initial data set to obtain a target intermediate data set; determining the attribute feature distribution information of the target intermediate data set and the logistics paths in each other initial data set under a preset attribute, where the other initial data sets are the initial data sets except the target initial data set; determining the feature distribution distance between the target intermediate data set and each other initial data set according to the attribute feature distribution information of the target intermediate data set and each other initial data set; and performing weighted averaging on the feature distribution distances between the target intermediate data set and each other initial data set to obtain the feature distribution difference coefficient of the target intermediate data set after adding the target remaining logistics path to the target initial data set.

[0086] Among them, the preset attribute can be an attribute in the logistics scenario, such as actual vehicle count, mileage, actual volume, actual cost, actual load rate, etc.

[0087] Among them, the feature distribution distance can be, for example, Wassertein Distance (a feature distribution distance that measures the data distribution in two data sets).

[0088] Further, step S207 may further include: adding the target remaining logistics path to the initial data set corresponding to the smallest feature distribution difference coefficient.

[0089] That is, when the current sample volume balance coefficient is less than or equal to the balance coefficient threshold, add the target remaining logistics path to the initial data set corresponding to the smallest feature distribution difference coefficient.

[0090] The smaller the feature distribution difference coefficient, the closer the feature distributions of the data sets are. Therefore, adding the target remaining logistics path to the initial data set corresponding to the smallest feature distribution difference coefficient can help the overall feature distribution of the initial data set to be closer.

[0091] Figure 4 The flowchart of the logistics data processing method according to an embodiment of the present disclosure is shown, as Figure 4 As shown, the logistics data processing method provided by the embodiment of the present disclosure may include the following steps.

[0092] Step 1: Route sorting.

[0093] In this step, considering that the shipping days of each route are different, routes with more shipping days have a greater impact on the splitting result (i.e., the result of determining the target data set), and routes with fewer shipping days have a smaller impact on the splitting result. Therefore, it is necessary to first calculate the shipping days of each route and sort each route in reverse order according to the shipping days.

[0094] Step 2: Group initialization (i.e., initialization of the initial data set).

[0095] In this step, 3k routes (where k is the number of groups) can be extracted from the routes that are shipped every day, and then 3 routes are assigned to each empty initial data set according to the ODT information of the routes by using the random hash splitting method. The purpose of doing this is to solve the cold start problem of the adaptive splitting algorithm.

[0096] Step 3: Adaptive splitting of the remaining routes (i.e., the target remaining logistics paths).

[0097] In this step, the remaining logistics path with the highest shipping frequency in the remaining logistics paths can be determined as the target remaining logistics path, and the current sample volume balance coefficient can be calculated. Then, according to the size relationship between the sample volume balance coefficient and the threshold (i.e., the balance coefficient threshold), the target remaining logistics path is split (i.e., determine the initial data set to which the target remaining logistics path should be added) in different cases.

[0098] Among them, the remaining route data can be sequentially taken out from the sorting result in Step 1 as the routes to be grouped, that is, as the target remaining logistics paths.

[0099] Specifically, if the balance coefficient (i.e., the sample size balance coefficient) is greater than the threshold value, the line to be grouped is placed into the group (initial data set) that can reduce the balance coefficient the most. The purpose is to make the sample sizes between groups as balanced as possible. Among them, the balance coefficient can be used to represent the imbalance degree of the sample quantities between groups. The larger the balance coefficient, the more unbalanced the sample quantities of each group are. The proportion of the sample size of each group is p k , and the expression of the balance coefficient is:

[0100]

[0101] If the balance coefficient is less than the threshold value, calculate the Wassertein Score (Wassertein score) of each group after the line to be grouped is placed into each group, and then place the line to be grouped into the group with the smallest Wassertein Score. Among them, the Wassertein Score is the weighted average of the Wassertein Distance between this group and other groups, and the weight of each Wassertein Distance can be the proportion of the sample size in other groups. Among them, the Wassertein Distance represents the distance of the feature distribution between two groups. The larger the value of the Wassertein Distance, the lower the similarity of the feature distributions of the two groups, and the more unbalanced the grouping. The purpose of weighting by the proportion of the sample size is to balance the sample sizes of each group as much as possible while considering the balance of the feature distribution. The formula for the Wassertein Distance is:

[0102]

[0103] Figure 5 shows a schematic diagram of the feature distribution result obtained in the logistics data processing method according to an embodiment of the present disclosure, as Figure 5 shown, it is a comparative schematic diagram of the probability density distribution diagrams of the features of the target data sets in the case of k = 2 (i.e., the number of target data sets is 2).

[0104] Among them, the left figure is the probability density distribution diagrams of target data set 1 and target data set 2 when the feature is the actual volume. Among them, the abscissa is the value of the feature of the actual volume, and the ordinate is the magnitude of the probability density corresponding to this value. The right figure is the probability density distribution diagrams of target data set 1 and target data set 2 when the feature is the actual load rate. Among them, the abscissa is the value of the feature of the actual load rate, and the ordinate is the magnitude of the probability density corresponding to this value.

[0105] According to Figure 5It can be seen that the probability density distributions of the target data set 1 and the target data set 2 are very close under two different characteristics, namely the actual volume and the actual load rate, indicating that the data distributions of these two target data sets are balanced and all characteristics are very close.

[0106] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0107] Figure 6 The block diagram of a logistics data processing device 600 according to an embodiment of the present disclosure is shown; as Figure 6 shown, it includes: an acquisition module 601, configured to acquire multiple logistics paths and the shipping frequencies of each logistics path; an initial data set division module 602, configured to determine target logistics paths from the multiple logistics paths according to the shipping frequencies of each logistics path, and divide the target logistics paths into a first number of initial data sets; a calculation module 603, configured to calculate, for each remaining logistics path except the target logistics paths, the data set balance coefficient after adding the remaining logistics path to each initial data set in sequence; a target data set determination module 604, configured to add each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient to obtain a first number of target data sets.

[0108] Through the logistics data processing device provided by the present disclosure, the target logistics paths can be selected first according to the shipping frequencies of multiple logistics paths, then the selected target logistics paths can be divided into a first number of initial data sets. Next, the data set balance coefficient after adding each remaining logistics path to each initial data set can be calculated for each remaining logistics path, and then each remaining logistics path can be automatically and adaptively split based on the data set balance coefficient, and the balance between all initial data sets after splitting can be ensured, so as to obtain a first number of target data sets with balanced data distributions. It can be seen that, on the one hand, this solution selects target logistics paths according to the shipping frequencies of multiple logistics paths and generates a first number of initial data sets based on this, and can realize an initialization method of the initial data set with the shipping frequency as a consideration factor, providing a cold start solution for the logistics data processing method provided by the present disclosure; on the other hand, this solution can automatically process each remaining logistics path based on the data set balance coefficient, and adaptively and evenly split each remaining logistics path. Therefore, all logistics paths can be processed while ensuring the overall balance of the initial data sets, and target data sets with balanced data distributions can be obtained. Furthermore, the target data sets with balanced data distributions can be further applied to the logistics production environment to provide a data basis for various logistics production requirements (such as strategy testing requirements).

[0109] In some embodiments, the dataset balance coefficient includes a sample size balance coefficient; wherein, for each remaining logistics path other than the target logistics path, the calculation module 603 sequentially calculates the dataset balance coefficient after adding the remaining logistics path to each initial dataset, including: determining the target remaining logistics path as the remaining logistics path with the highest shipping frequency in the remaining logistics paths; determining the current sample size balance coefficient according to the logistics paths in all initial datasets; if the current sample size balance coefficient is greater than the balance coefficient threshold, calculating the sample size balance coefficient after adding the target remaining logistics path to each initial dataset; wherein, the target dataset determination module 604 adds each remaining logistics path to one of the initial datasets according to the dataset balance coefficient, including: adding the target remaining logistics path to the initial dataset corresponding to the smallest sample size balance coefficient.

[0110] In some embodiments, the dataset balance coefficient further includes a feature distribution difference coefficient; wherein, for each remaining logistics path other than the target logistics path, the calculation module 603 sequentially calculates the dataset balance coefficient after adding the remaining logistics path to each initial dataset, further including: if the current sample size balance coefficient is less than or equal to the balance coefficient threshold, calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial dataset; wherein, the target dataset determination module 604 adds each remaining logistics path to one of the initial datasets according to the dataset balance coefficient, further including: adding the target remaining logistics path to the initial dataset corresponding to the smallest feature distribution difference coefficient.

[0111] In some embodiments, the calculation module 603 determines the current sample size balance coefficient according to the logistics paths in all initial datasets, including: determining the shipping days of the logistics paths in each initial dataset within the historical period according to the shipping frequencies of the logistics paths in each initial dataset; determining the sample size included in each initial dataset according to the shipping days of the logistics paths in each initial dataset; and determining the current sample size balance coefficient according to the sample sizes included in all initial datasets.

[0112] In some embodiments, the computing module 603 calculates the coefficient of difference in feature distribution after adding the target remaining logistics path to each initial data set, including: for each target initial data set, adding the target remaining logistics path to the target initial data set to obtain a target intermediate data set; determining the attribute feature distribution information of the target intermediate data set and each other initial data set according to the attribute values of the logistics paths in the target intermediate data set and each other initial data set under a preset attribute; the other initial data sets are the initial data sets except the target initial data set; determining the feature distribution distance between the target intermediate data set and each other initial data set according to the attribute feature distribution information of the target intermediate data set and each other initial data set; performing a weighted average on the feature distribution distances between the target intermediate data set and each other initial data set to obtain the coefficient of difference in feature distribution of the target intermediate data set.

[0113] In some embodiments, the initial data set partitioning module 602 determines a target logistics path from multiple logistics paths according to the shipping frequency of each logistics path, and partitions the target logistics path into a first number of initial data sets, including: sorting the multiple logistics paths in descending order according to the shipping frequency of each logistics path, and determining the first second number of logistics paths as the target logistics path; wherein, the second number is an integer multiple of the first number; evenly and randomly partitioning the target logistics path into a first number of initial data sets; wherein, the number of target logistics paths included in each initial data set is the same.

[0114] In some embodiments, the logistics data processing device further includes a testing module 605, and the testing module 605 is configured to: obtain a first number of logistics strategies to be tested; match the first number of target data sets with the first number of logistics strategies to be tested one by one to obtain a first number of test combinations; apply the logistics strategy to be tested in each test combination to the target data set to obtain the test effect corresponding to the test combination; use the logistics strategy to be tested in the test combination with the best test effect as the target logistics strategy, so as to apply the target logistics strategy to each logistics path.

[0115] Figure 6 For other contents of the embodiment, reference may be made to the above other embodiments.

[0116] Those skilled in the art to which the present invention pertains can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0117] Figure 7The structural block diagram of a computer device for processing logistics data in an embodiment of the present disclosure is shown. It should be noted that the illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0118] The following refers to Figure 7 to describe the electronic device 700 according to this embodiment of the present invention. Figure 7 The illustrated electronic device 700 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0119] As Figure 7 shown, the electronic device 700 is presented in the form of a general computing device. The components of the electronic device 700 may include, but are not limited to: at least one of the above-mentioned processing units 710, at least one of the above-mentioned storage units 720, and a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710).

[0120] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 710 can execute as Figure 2 shown in

[0121] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202, and may further include a read-only storage unit (ROM) 7203.

[0122] The storage unit 720 may further include a program / utility 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0123] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0124] The electronic device 700 can also communicate with one or more external devices 800 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0125] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0126] The program product for implementing the above method according to an embodiment of the present invention can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0127] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0128] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0129] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0130] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0131] It should be noted that although several modules or units of a device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.

[0132] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0133] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0134] According to one aspect of the present disclosure, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the methods provided in the various alternative implementations of the above embodiments.

[0135] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A logistics data processing method, characterized in that, Including: Obtaining a plurality of logistics paths and the shipping frequencies of each logistics path; Determining a target logistics path from the plurality of logistics paths according to the shipping frequencies of each logistics path, and dividing the target logistics path into a first number of initial data sets; For each remaining logistics path except the target logistics path, successively calculating the data set balance coefficient after adding the remaining logistics path to each initial data set; Adding each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient to obtain a first number of target data sets.

2. The method according to claim 1, characterized in that The data set balance coefficient includes a sample size balance coefficient; Among them, for each remaining logistics path except the target logistics path, successively calculating the data set balance coefficient after adding the remaining logistics path to each initial data set includes: Determining the target remaining logistics path as the remaining logistics path with the largest shipping frequency in the remaining logistics paths; Determining the current sample size balance coefficient according to the logistics paths in all initial data sets; If the current sample size balance coefficient is greater than the balance coefficient threshold, calculating the sample size balance coefficient after adding the target remaining logistics path to each initial data set; Among them, adding each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient includes: Adding the target remaining logistics path to the initial data set corresponding to the smallest sample size balance coefficient.

3. The method according to claim 2, characterized in that, The data set balance coefficient further includes a feature distribution difference coefficient; Among them, for each remaining logistics path except the target logistics path, successively calculating the data set balance coefficient after adding the remaining logistics path to each initial data set further includes: If the current sample size balance coefficient is less than or equal to the balance coefficient threshold, calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial data set; Among them, adding each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient further includes: Adding the target remaining logistics path to the initial data set corresponding to the smallest feature distribution difference coefficient.

4. The method according to claim 2, wherein Determining the current sample size balance coefficient according to the logistics paths in all initial data sets includes: Determining the shipping days of the logistics paths in each initial data set within a historical period according to the shipping frequencies of the logistics paths in each initial data set; Determining the sample size included in each initial data set according to the shipping days of the logistics paths in each initial data set; Determining the current sample size balance coefficient according to the sample sizes included in all initial data sets.

5. The method according to claim 3, characterized in that, Calculating the feature distribution difference coefficient after adding the target remaining logistics path to each initial data set includes: For each target initial data set, adding the target remaining logistics path to the target initial data set to obtain a target intermediate data set; Determine the attribute feature distribution information of the target intermediate data set and each other initial data set according to the attribute values of the logistics paths in the target intermediate data set and each other initial data set under preset attributes; the other initial data sets are the initial data sets except the target initial data set; Determine the feature distribution distance between the target intermediate data set and each other initial data set according to the attribute feature distribution information of the target intermediate data set and each other initial data set; Perform weighted averaging on the feature distribution distances between the target intermediate data set and each other initial data set to obtain the feature distribution difference coefficient of the target intermediate data set.

6. The method according to claim 1, wherein Determine the target logistics path from the multiple logistics paths according to the shipping frequencies of each logistics path, and divide the target logistics path into the first number of initial data sets, including: Sort the multiple logistics paths in descending order according to the shipping frequencies of each logistics path, and determine the first second number of logistics paths as the target logistics path; wherein, the second number is an integer multiple of the first number; Divide the target logistics path evenly and randomly into the first number of initial data sets; wherein, the number of target logistics paths included in each initial data set is the same.

7. The method according to any one of claims 1-6, characterized in that The method further includes: Obtain the first number of logistics strategies to be tested; Match the first number of target data sets with the first number of logistics strategies to be tested one by one to obtain the first number of test combinations; Apply the logistics strategy to be tested in each test combination to the target data set to obtain the test effect corresponding to the test combination; Use the logistics strategy to be tested in the test combination with the best test effect as the target logistics strategy to apply the target logistics strategy to each logistics path.

8. A logistics data processing device, characterized in that, Including: An acquisition module, configured to acquire a plurality of logistics paths and the shipping frequency of each logistics path; An initial data set division module, configured to determine a target logistics path from the plurality of logistics paths according to the shipping frequencies of each logistics path, and divide the target logistics path into the first number of initial data sets; A calculation module, configured to calculate, for each remaining logistics path except the target logistics path, the data set balance coefficient after adding the remaining logistics path to each initial data set in sequence; A target data set determination module, configured to add each remaining logistics path to one of the initial data sets respectively according to the data set balance coefficient to obtain the first number of target data sets.

9. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the logistics data processing method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the logistics data processing method according to any one of claims 1 to 7.