Parking site mining method and device and storage medium
Through density-based clustering algorithm and time series feature analysis, combined with similarity calculation, the station points can be accurately identified, solving the problem of inaccurate station point identification and optimizing logistics operations.
Patent Information
- Application Number
- CN202410501423.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-10-24
AI Technical Summary
In the existing technology, the identification of the on-site points is inaccurate, which makes it impossible to reduce the difficulty of the collection operation and optimize the operating costs.
By obtaining the geographic location of the shipping addresses of non-resident customers, clustering is performed using a density-based clustering algorithm, and combining the number of waybills, seasonal characteristics, and future time series characteristics, candidate shipping points are determined. The shipping points are then accurately identified through similarity calculation and feature screening.
It improves the accuracy of on-site point identification, ensures that the service scope and waybill volume are appropriate, and reduces the difficulty of collection operations and operating costs.
Smart Images

Figure CN120833099A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of logistics and computers, and particularly relates to a method and device for mining a field point and a storage medium. BACKGROUND
[0002] In the logistics collection link, for a customer (such as a merchant) with a large number of shipments, a field point is set up at the shipping address of the customer, and a field staff is configured to be responsible for the shipment business of the field customer, such as shipment guarantee and follow-up service, resource allocation, and shipment order exception handling, which helps to improve customer experience and satisfaction, and also helps to reduce the difficulty of collection operation and reduce the operating cost of the logistics company.
[0003] In some related technologies, the number of shipments collected by a certain customer (such as a merchant) in a specified historical time period is counted, and if the number of shipments reaches a certain threshold, the customer is recommended as a field customer, and a field point is set up for the field customer. SUMMARY
[0004] Identifying a field point in the dimension of a customer (such as a merchant) may have the problem of inaccurate identification of a field point. For example, the number of shipments of merchant A in a certain month reaches the threshold of a field point, and according to the related technology, a field point should be set up for merchant A, but the actual situation is that merchant A has multiple shipping addresses, the shipping addresses are far apart, and the number of shipments of each shipping address does not reach the threshold, so the difficulty of collection operation will not be reduced by setting up a field point, and therefore it is not suitable to set up a field point for merchant A. For another example, if the shipping addresses of multiple merchants are adjacent and can be regarded as the same shipping address, and the total number of shipments of these merchants reaches the threshold of a field point, it is appropriate to set up a field point for these merchants jointly, but according to the related technology, since the number of shipments of each merchant does not reach the threshold of a field point, a field point will not be set up for these merchants.
[0005] Therefore, some embodiments of the present disclosure provide a method for mining a field point, which comprises: obtaining the geographic positions of the shipping addresses of each non-field customer; taking the geographic position of each shipping address of each non-field customer as a sample point to construct a sample point set; clustering each sample point in the sample point set to obtain each sample point cluster; determining a candidate field point according to each sample point cluster; and determining a field point according to the candidate field point.
[0006] According to the clustering of the geographic positions of the shipping addresses of each non-field customer, a candidate field point is determined, and then a field point is determined, thereby improving the accuracy of identification of a field point.
[0007] In some embodiments, clustering each sample point in the sample point set to obtain each sample point cluster comprises: clustering each sample point in the sample point set by using a density-based clustering algorithm, and imposing a constraint on the clustering result, so that sample points of the same type and in the same site are aggregated into one sample point cluster.
[0008] The density-based clustering algorithm can automatically detect clusters of arbitrary shape and quantity, and is suitable for implementing clustering of the geographic location of the sending address in the station point business scenario. In addition, since the cluster sizes obtained by the density-based clustering algorithm are different, and the clustering result is greatly affected by parameters, by imposing a constraint on the clustering result, the adaptability of the clustering result in the station point business scenario can be improved, and thus the accuracy of station point identification can be improved.
[0009] In some embodiments, determining the candidate station point according to each sample point cluster comprises: calculating a distance mean based on the distance between each sample point in each sample point cluster; and determining a sample point cluster with a distance mean less than a distance threshold as a candidate station point.
[0010] The distance mean between each sample point in the sample point cluster is used to measure the size of the sample point cluster, and by comparing with the distance threshold, the size of the sample point cluster, i.e. the service range of the station point, is limited to an appropriate size.
[0011] In some embodiments, determining the station point according to the candidate station point comprises: determining the station point according to the order quantity of the candidate station point.
[0012] After determining the appropriate candidate station point, the candidate station point with a qualified order quantity is determined as the station point in combination with the order quantity of the candidate station point, so that a station point with appropriate spatial dimensions and order quantity is found, and the accuracy of station point identification is further improved.
[0013] In some embodiments, determining the station point according to the order quantity of the candidate station point comprises: removing candidate station points with long-tail characteristics according to time series data of the order quantity of the candidate station point; and determining the station point according to the remaining candidate station points.
[0014] By removing candidate station points with increasingly smaller order quantities and retaining candidate station points with qualified order quantities, the accuracy of station point identification is further improved.
[0015] In some embodiments, determining the station point according to the candidate station point comprises: determining the station point according to one or more of seasonal characteristics and future time series characteristics of the order quantity of the candidate station point.
[0016] The seasonal characteristics of the shipment quantity of the candidate resident points, the future shipment quantity, and the like are comprehensively considered to find the resident points that meet the requirements of the seasonal characteristics of the shipment quantity and the future shipment quantity, and the accuracy of the identification of the resident points is further improved.
[0017] In some embodiments, the resident points are determined according to one or more of the seasonal characteristics of the shipment quantity of the candidate resident points and the future time sequence characteristics of the shipment quantity. The candidate resident points with seasonal fluctuations are removed according to the seasonal characteristics of the shipment quantity of the candidate resident points, and / or the candidate resident points with future shipment quantity lower than a threshold are removed according to the future time sequence characteristics of the shipment quantity of the candidate resident points. The resident points are determined according to the remaining candidate resident points.
[0018] The candidate resident points with seasonal fluctuations are removed, and the candidate resident points with future shipment quantity lower than a threshold are removed, according to the seasonal characteristics of the shipment quantity of the candidate resident points, the future time sequence characteristics of the shipment quantity, and the like, and the accuracy of the identification of the resident points is further improved.
[0019] In some embodiments, the method further includes determining the seasonal characteristics of the shipment quantity of the candidate resident points, including: obtaining each resident point cluster with different time sequence pattern shape categories by using a shape-based time sequence clustering algorithm according to the time sequence data of the shipment quantity of each candidate resident point; and determining the seasonal characteristics of each resident point cluster as the seasonal characteristics of the shipment quantity of the candidate resident points in the each resident point cluster.
[0020] In the resident point business, the number of resident points is large, and the seasonal analysis method for a single time sequence is difficult to apply, and the seasonal characteristics of a single resident point are not the focus of mining. After clustering each resident point by using a shape-based time sequence clustering algorithm, the seasonal characteristics of the resident points are labeled according to the seasonal characteristics of the large categories, so that the seasonal characteristics are more referential when the resident points are decided.
[0021] In some embodiments, the seasonal characteristics of each resident point cluster include one or more of the number of average annual shipment months, the proportion of average monthly shipment quantity to annual shipment quantity, and the number of peak values.
[0022] Thus, the seasonal characteristics affecting the decision of the resident points are determined.
[0023] In some embodiments, the method further includes determining the future time sequence characteristics of the shipment quantity of the candidate resident points, including: establishing a time sequence prediction algorithm of each candidate resident point according to the time sequence data of the shipment quantity of the candidate resident point, for predicting the future time sequence characteristics of the shipment quantity of the candidate resident point, wherein the time sequence prediction algorithm includes a growth trend item, a periodic item, a holiday effect, and an error item.
[0024] Based on a time series prediction algorithm, the future time sequence characteristics of the shipment volume of the candidate resident point are determined more accurately by comprehensively considering the growth trend, cycle, holiday effect, and error, and the resident point is determined more accurately.
[0025] In some embodiments, the future time sequence characteristics of the shipment volume of the candidate resident point include one or more of the shipment volume of the candidate resident point at a future time, the shipment frequency characteristics of the candidate resident point, and the future variation trend characteristics of the shipment volume of the candidate resident point.
[0026] Thus, the future time sequence characteristics of the shipment volume affecting the resident point decision are determined.
[0027] In some embodiments, determining the resident point according to the candidate resident point includes determining the resident point according to the similarity of the candidate resident point to the existing resident point set.
[0028] Thus, the newly determined resident point is more suitable for the actual operation, and the accuracy of resident point identification is improved.
[0029] In some embodiments, the method further includes determining the similarity of the candidate resident point to the existing resident point set, including: calculating the similarity of each candidate resident point to each existing resident point in the existing resident point set; determining the weight of each existing resident point as the proportion of the shipment volume of the existing resident point to the total shipment volume of the existing resident point set; and performing weighted calculation according to the similarity of each candidate resident point to each existing resident point in the existing resident point set and the weight of each existing resident point to obtain a weighted similarity as the similarity of each candidate resident point to the existing resident point set.
[0030] The weight of each existing resident point is determined by the proportion information of the shipment volume of the existing resident point, and the similarity of the candidate resident point to the existing resident point set is determined more accurately by weighted calculation, and thus the newly determined resident point is more suitable for the actual operation, and the accuracy of resident point identification is improved.
[0031] In some embodiments, calculating the similarity of each candidate resident point to each existing resident point in the existing resident point set includes calculating the similarity of the feature vector of each candidate resident point to the feature vector of each existing resident point in the existing resident point set as the similarity of each candidate resident point to each existing resident point in the existing resident point set, wherein the plurality of features are determined according to the plurality of features affecting the resident point decision.
[0032] The similarity of the candidate resident point to the existing resident point set is determined more accurately by calculating the similarity of the feature vector of the plurality of features affecting the resident point decision, and thus the newly determined resident point is more suitable for the actual operation, and the accuracy of resident point identification is improved.
[0033] In some embodiments, the method further comprises: screening, by the tree model, a plurality of initial features affecting the station point decision; and constructing a feature vector of the plurality of features based on the plurality of features meeting the requirement in terms of importance.
[0034] Screening the plurality of important features affecting the station point decision and constructing the feature vector of the plurality of features make the similarity between the candidate station point and the existing station point set calculated therefrom more accurate, and further make the newly determined station point according to the similarity more suitable for the actual operation, thereby improving the accuracy of station point identification.
[0035] In some embodiments, the plurality of features include one or more of a time sequence feature of the order quantity, a contracted discount feature, a cargo side quantity feature, a human efficiency feature, a user category feature, and a express product type feature.
[0036] Thus, the plurality of features affecting the station point decision include not only the order quantity related features but also other related features, thereby improving the accuracy of station point identification.
[0037] Some embodiments of the present disclosure provide a station point mining apparatus, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute a station point mining method based on instructions stored in the memory.
[0038] Some embodiments of the present disclosure provide a station point mining apparatus, comprising: a module for executing a station point mining method.
[0039] Some embodiments of the present disclosure provide a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement steps of a station point mining method.
[0040] Some embodiments of the present disclosure provide a computer program product comprising a computer program, the computer program being executed by a processor to implement steps of a station point mining method. BRIEF DESCRIPTION OF DRAWINGS
[0041] The drawings needed to be used in the following embodiments or related technical descriptions will be briefly introduced. According to the following detailed description with reference to the drawings, the present disclosure can be more clearly understood.
[0042] Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0043] Figure 1 A flowchart of a station point mining method of some embodiments of the present disclosure is shown.
[0044] Figure 2 A flowchart showing a method for mining a stationary point according to some embodiments of the present disclosure.
[0045] Figure 3 A flowchart showing a method for mining a stationary point according to some embodiments of the present disclosure.
[0046] In Figure 4 (a) shows the meaning of the parameters of DBSCAN, for example, the sample neighborhood radius ∈ is 1, and the minimum number of sample points in the sample neighborhood radius MinPts is 4; (b) shows three examples of sample points of DBSCAN; (c) shows four examples of sample point relationships of DBSCAN; (d) shows an example of the clustering step of DBSCAN.
[0047] In Figure 5 (a) represents a time series data A, (b) represents another time series data B, (c) represents the similarity of time series data A and time series data B calculated according to K-means, (d) represents the similarity of time series data A and time series data B calculated according to K-shape algorithm, and (e) represents an example of the clustering result of K-shape algorithm.
[0048] In Figure 6 (a) represents a Prophet algorithm prediction result schematic diagram, and (b) represents the effect result schematic diagram of each decomposition term of the Prophet algorithm.
[0049] Figure 7 A structural schematic diagram of a device for mining a stationary point according to some embodiments of the present disclosure is shown.
[0050] Figure 8 A structural schematic diagram of a device for mining a stationary point according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0051] It should be noted that, unless otherwise specified, the relative arrangement, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0052] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they represent the inevitable logical sequence between them.
[0053] It should also be understood that in the embodiments of the present disclosure, "a plurality of" can mean two or more, and "at least one" can mean one, two or more.
[0054] It should also be appreciated that for any relative term mentioned in the embodiments of the present disclosure, one or more can be understood in general, without explicit limitation or in the context of the opposite indication.
[0055] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the front and rear associated objects.
[0056] It should also be understood that the description of the embodiments of the present disclosure emphasizes the differences between the embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0057] At the same time, it should be understood that, for the convenience of description, the size of each part shown in the drawings is not drawn according to the actual proportion relationship.
[0058] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the disclosure, its application or uses.
[0059] The techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods, and devices should be considered as part of the specification.
[0060] It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0061] In addition, in order to avoid obscuring the present disclosure due to unnecessary details, only the processing steps and / or device structures closely related to the scheme according to the present disclosure are shown in the drawings, and other details not closely related to the present disclosure are omitted. It should also be noted that similar reference numerals and letters in the drawings indicate similar items, and therefore once an item is defined in one drawing, it does not need to be discussed again for subsequent drawings.
[0062] For the foregoing, the customer (such as a merchant) dimension is used to identify the site, and there can be a problem of inaccurate site identification. The site mining method according to the embodiments of the present disclosure is proposed to improve the accuracy of site identification.
[0063] Figure 1 The flowchart of the site mining method according to some embodiments of the present disclosure is shown.
[0064] As Figure 1As shown, the embodiment of the method for mining the resident point comprises: step 110, obtaining the geographic positions of the sending addresses of each non-resident customer; step 120, taking the geographic position of each sending address of each non-resident customer as a sample point, and constructing a sample point set; step 130, clustering each sample point in the sample point set to obtain each sample point cluster; step 140, determining a candidate resident point according to each sample point cluster; and step 150, determining a resident point according to the candidate resident point.
[0065] According to the clustering of the geographic positions of the sending addresses of each non-resident customer, the candidate resident point is determined, and then the resident point is determined, thereby improving the accuracy of the identification of the resident point.
[0066] In step 110, the non-resident customer refers to a customer who has not yet set up a resident point, and the customer may include, for example, a merchant who frequently mails goods. The sending address of the customer refers to the address where the customer actually ships goods. Each customer may have one or more sending addresses. The geographic positions of the multiple sending addresses of the customer are dispersed, and some sending addresses may be relatively close, and some sending addresses may be relatively far apart. The geographic position of the sending address may be, for example, the latitude and longitude information of the sending address.
[0067] In step 120, the geographic position of each sending address of each non-resident customer is taken as a sample point, and if a non-resident customer has multiple sending addresses, the geographic positions of the multiple sending addresses of the non-resident customer correspond to multiple sample points.
[0068] In step 130, clustering each sample point in the sample point set to obtain each sample point cluster comprises: using a density-based clustering algorithm to cluster each sample point in the sample point set, and imposing a constraint on the clustering result, so that the clustering result is aggregated into a sample point cluster for sample points of the same type and in the same site. The density-based clustering algorithm may be, for example, a DBSCAN algorithm.
[0069] Using the density-based clustering algorithm can automatically detect clusters of arbitrary shape and quantity, which is suitable for implementing the clustering of the geographic positions of the sending addresses in the resident point business scenario. In addition, because the cluster sizes obtained by the density-based clustering algorithm are different, and the clustering result is greatly affected by the parameters, by imposing a constraint on the clustering result, the adaptability of the clustering result in the resident point business scenario can be improved, and then the accuracy of the identification of the resident point can be improved.
[0070] In step 140, if no restriction is imposed, one sample point cluster may correspond to one candidate resident point. According to the actual situation of the resident point, a restriction is imposed, and then determining a candidate resident point according to each sample point cluster comprises: calculating the distance mean based on the distance between each sample point in each sample point cluster; and determining a sample point cluster with a distance mean less than a distance threshold as a candidate resident point.
[0071] The size of the sample point cluster is measured by the mean distance between each sample point in the sample point cluster, and by comparing it with the distance threshold, the size of the sample point cluster, that is, the service range of the station point, is limited to an appropriate size.
[0072] In step 150, a station point represents a virtual location range suitable for configuring station personnel and a customer list within the virtual location range. The station point can be simply referred to as a station point. There are many different implementation methods for determining the station point based on the candidate station points. These different implementation methods can be used alone or in combination. Figure 2 describe.
[0073] Figure 2 A schematic flow chart illustrating a station point excavation method according to some embodiments of the present disclosure is provided. Figure 2 relatively Figure 1 , adding descriptions of various different implementation methods of step 150. Figure 2 As shown, the site excavation method of this embodiment includes steps 110-140, and further includes one or more of steps 150A, 150B, 150C, and 150D.
[0074] In step 150A, determining a candidate station location includes determining the station location based on the number of waybills at the candidate station location. The number of waybills at the candidate station location is the total number of waybills for shipping addresses of customers not present at the candidate station location. Examples of the number of waybills at the candidate station location include, but are not limited to, historical number of waybills and predicted future number of waybills.
[0075] After determining the appropriate candidate stationing points, the candidate stationing points with the required number of waybills will be determined as the stationing points based on the number of waybills at the candidate stationing points, thereby finding the stationing points with appropriate spatial dimensions and number of waybills, and further improving the accuracy of stationing point identification.
[0076] In step 150B, determining the candidate station points includes: removing candidate station points with long-tail characteristics based on the time series data of the waybill volume of the candidate station points; and determining the station points based on the remaining candidate station points. The time series data of the waybill volume refers to the sequence of waybill volumes at different times, and the granularity of this time can be set, such as daily or monthly. Candidate station points with long-tail characteristics are candidate station points whose waybill volume decreases over time. The graphical representation of these decreasing waybill volumes resembles a "long tail."
[0077] By removing candidate depots with decreasing number of waybills and retaining candidate depots with qualified number of waybills, the accuracy of depot identification can be further improved.
[0078] In step 150C, determining the resident point according to the candidate resident point includes: determining the resident point according to one or more of the seasonal characteristics, the future timing characteristics of the shipment quantity of the candidate resident point.
[0079] The seasonal characteristics, the future timing characteristics, etc. of the shipment quantity of the candidate resident point are comprehensively considered to find the resident point whose seasonal characteristics and future shipment quantity meet the requirements, and the accuracy of the resident point identification is further improved.
[0080] In some embodiments, determining the resident point according to one or more of the seasonal characteristics, the future timing characteristics of the shipment quantity of the candidate resident point includes: removing the candidate resident point with seasonal fluctuations according to the seasonal characteristics of the shipment quantity of the candidate resident point, and / or removing the candidate resident point with future shipment quantity lower than a threshold according to the future timing characteristics of the shipment quantity of the candidate resident point; determining the resident point according to the remaining candidate resident points.
[0081] The candidate resident points with seasonal fluctuations are removed according to the seasonal characteristics of the shipment quantity of the candidate resident point, and the candidate resident points with future shipment quantity lower than a threshold are removed according to the future timing characteristics of the shipment quantity of the candidate resident point, and the accuracy of the resident point identification is further improved.
[0082] Determining the seasonal characteristics of the shipment quantity of the candidate resident point includes: obtaining each resident point cluster with different timing pattern shape categories by using a shape-based timing clustering algorithm according to the timing data of the shipment quantity of each candidate resident point; determining the seasonal characteristics of each resident point cluster as the seasonal characteristics of the shipment quantity of the candidate resident point in the each resident point cluster. The shape-based timing clustering algorithm is, for example, a K-shape algorithm.
[0083] In the resident point business, the resident point quantity is large, and the seasonal analysis method for a single time series is difficult to apply, and the seasonal characteristics of a single resident point are not the focus of mining. After clustering each resident point by using a shape-based timing clustering algorithm, the seasonal characteristics of the resident point are labeled according to the seasonal characteristics of the large category, so that the seasonal characteristics are more referential when the resident point is decided.
[0084] The seasonal characteristics of each resident point cluster include, but are not limited to, one or more of: the average number of delivery months per year of each resident point cluster, the average monthly shipment quantity proportion of the shipment quantity of the current year, and the number of peak values. Thus, the seasonal characteristics affecting the resident point decision are determined.
[0085] The future time series feature of the shipment volume of the candidate resident point includes, for example, establishing a time series prediction algorithm, such as a Prophet algorithm, for each candidate resident point according to the time series data of the shipment volume of each candidate resident point, to predict the future time series feature of the shipment volume of each candidate resident point, wherein the time series prediction algorithm includes a growth trend item, a periodicity item, a holiday effect, and an error item.
[0086] Based on the time series prediction algorithm, such as the Prophet algorithm, the growth trend, the periodicity, the holiday effect, and the error are comprehensively considered to more accurately determine the future time series feature of the shipment volume of the candidate resident point, and thus more accurately determine the resident point.
[0087] The future time series feature of the shipment volume of the candidate resident point includes, for example, but is not limited to, one or more of the shipment volume of the candidate resident point at a future time, the shipment frequency feature of the candidate resident point, and the future variation trend feature of the shipment volume of the candidate resident point. Thus, the future time series feature of the shipment volume affecting the resident point decision is determined.
[0088] In step 150D, determining the resident point according to the candidate resident point includes determining the resident point according to the similarity of the candidate resident point to the existing resident point set.
[0089] Thus, the newly determined resident point is more suitable for the actual operation, and the accuracy of resident point identification is improved.
[0090] The similarity of the candidate resident point to the existing resident point set is determined, for example, by calculating the similarity of each candidate resident point to each existing resident point in the existing resident point set, determining the weight of each existing resident point as the proportion of the shipment volume of the existing resident point to the total shipment volume of the existing resident point set, and performing weighted calculation according to the similarity of each candidate resident point to each existing resident point in the existing resident point set and the weight of each existing resident point to obtain a weighted similarity as the similarity of each candidate resident point to the existing resident point set.
[0091] The weight of the existing resident point is determined by the proportion information of the shipment volume of the existing resident point, and the similarity of the candidate resident point to the existing resident point set is more accurately determined by weighted calculation, and thus the newly determined resident point is more suitable for the actual operation, and the accuracy of resident point identification is improved.
[0092] The similarity of each candidate resident point to each existing resident point in the existing resident point set is calculated, for example, by calculating the similarity of the feature vector of each candidate resident point to the feature vector of each existing resident point in the existing resident point set, as the similarity of each candidate resident point to each existing resident point in the existing resident point set, wherein the plurality of features are determined according to the plurality of features affecting the resident point decision.
[0093] By calculating the similarity of the feature vector of the plurality of features affecting the decision of the resident point, the similarity of the candidate resident point to the existing resident point set is more accurately determined, and thus the newly determined resident point is more suitable for the actual operation, and the accuracy of the resident point identification is improved.
[0094] Among them, the plurality of initial features affecting the decision of the resident point are screened through the tree model; and the feature vector of the plurality of features is constructed based on the plurality of features whose importance degree meets the requirement.
[0095] The plurality of important features affecting the decision of the resident point are screened out, and the feature vector of the plurality of features is constructed, so that the similarity of the candidate resident point to the existing resident point set calculated therefrom is more accurate, and thus the newly determined resident point according to the similarity is more suitable for the actual operation, and the accuracy of the resident point identification is improved.
[0096] Among them, the plurality of features include one or more of the following: time sequence feature of waybill quantity, discount feature of signing, feature of cargo side quantity, feature of human efficiency, feature of user category, feature of express product type.
[0097] Therefore, the plurality of features affecting the decision of the resident point not only include features related to the waybill quantity, but also include other related features, thereby improving the accuracy of the resident point identification.
[0098] As described above, there are a plurality of different implementation methods for determining the resident point according to the candidate resident point in step 150, and these different implementation methods can be used alone or jointly. The following describes an exemplary method of joint use, and those skilled in the art can understand that the joint use is not limited to this exemplary method, and there can be other joint use methods. Figure 3
[0099] Figure 3 A flowchart of a resident point mining method of some embodiments of the present disclosure is shown. Figure 3 In comparison Figure 2 , an exemplary method of joint use of a plurality of different implementation methods for determining the resident point according to the candidate resident point is described, in which the customer is described as an example of a merchant.
[0100] Figure 3 The main process of the embodiment includes: filtering out non-resident merchant waybills for analysis based on waybill data, then generating resident points based on geographic information data using a density clustering algorithm, then constructing waybill quantity time sequence data based on resident point dimensions, then evaluating resident point waybill quantity variation based on time sequence clustering and prediction algorithm, and finally recommending resident points and corresponding merchants in combination with the results of collaborative filtering algorithm.
[0101] Figure 3 The resident point mining method of the embodiment mainly includes steps 310, 320, 330 and 340. The specific implementation process of each step and the related principles are described as follows.
[0102] Step 310, basic data construction.
[0103] In step 311, all the pickup orders and the corresponding feature field data of the merchant in a specified historical period (such as the past two years) are captured. The feature fields include, for example, merchant number and name, pickup site and type, pickup person and post, mailing address and corresponding geographic location information, express product type, consignment information, etc. In step 312, the orders are labeled based on the feature fields, such as “resident merchant” or “non-resident merchant”, to facilitate the screening of non-resident merchant orders for analysis. The pickup merchant order details table obtained is shown in Table 1 for example.
[0104] Table 1: Pickup merchant order details table
[0105]
[0106] Step 320, merchant address clustering.
[0107] The commonly used clustering algorithms can be divided into three categories: hierarchical method, partition method and density-based method. The density-based method can automatically detect clusters of arbitrary shape and number, and is suitable for implementing merchant address clustering in the resident business scenario, thereby generating the resident points to be analyzed (i.e. candidate resident points). The density clustering of the embodiment is implemented based on the DBSCAN algorithm for example.
[0108] DBSCAN algorithm describes the sample density by two algorithm parameters (v, MinPts), where ∈ is the sample neighborhood radius and MinPts is the minimum number of sample points within the sample neighborhood radius. Based on ∈ and MinPts, sample points are divided into three types: ① core point: if the number of sample points within the neighborhood radius ∈ of a sample point is greater than or equal to MinPts, the sample point is a core point, ② boundary point: a sample point that is not a core point but is within the ∈ neighborhood range of a core point, ③ noise point: a sample point that is neither a core point nor a boundary point. There are four relationships between sample points: ① density direct: if P is a core point and Q is within the ∈ neighborhood of P, then P is density direct to Q, the density direct relationship is not symmetric, ② density reachable: if there are core points P1, P2, P3, …, Pn, and P1 is density direct to P2, P2 is density direct to P3, …, P(n-1) is density direct to Pn, and Pn is density direct to Q, then P1 is density reachable to Q, the density reachable relationship is not symmetric, ③ density connected: for sample points P and Q, if there is a core point S such that S is density reachable to P and Q, then P and Q are density connected, the density connected relationship is symmetric, and the two points that are density connected belong to the same cluster, ④ non-density connected: if two points do not belong to the density connected relationship, then the two points are non-density connected, and the two points that are non-density connected belong to different clusters or there is a noise point among them.
[0109] The clustering steps of the DBSCAN algorithm mainly include the following two steps. ① Finding core points to form temporary cluster: scan all sample points, find core points and include them in the core point list, and form corresponding temporary clusters for the density direct points of the core points; ② merging connected temporary clusters to obtain final clusters: for a temporary cluster, check whether the points in the temporary cluster are core points, if so, merge the temporary cluster corresponding to the point with the current temporary cluster to obtain a new temporary cluster, repeat the operation until the points that are density connected with the sample points in the current temporary cluster are all in the temporary cluster, then the temporary cluster is a final cluster. The final generated clustering result includes clusters and noise points, and the clusters contain core points and boundary points.
[0110] In Figure 4 (a) shows the meaning of the parameters of DBSCAN, for example, the sample neighborhood radius ∈ is 1 and the minimum number of sample points within the sample neighborhood radius MinPts is 4; (b) shows an example of the three types of sample points of DBSCAN; (c) shows an example of the four relationships between sample points of DBSCAN; (d) shows an example of the clustering steps of DBSCAN.
[0111] Since the cluster size obtained by DBSCAN clustering is different, and the clustering result is greatly affected by the parameters, the clustering result needs to be subjected to constraint conditions, and the clustering results of DBSCAN model at multiple different parameters are integrated, and finally the analyzed resident points are output.
[0112] The process of clustering merchant addresses based on the DBSCAN algorithm is described below.
[0113] Step 321: Filter out the non-resident merchant shipment manifest to be analyzed from the pickup merchant shipment manifest, and arrange to obtain the non-resident merchant address information table, i.e. the merchant address information table to be analyzed, for example, as shown in Table 2.
[0114] Table 2 merchant address information example
[0115]
[0116]
[0117] Step 322: Density clustering. The longitude and latitude of the sender address of each merchant are taken as a sample point, and a sample point set is obtained: {(lng i ,lat i )|i=1,2,3,…,n}, lng i ,lat i respectively represent the longitude and latitude of sample point i, and DBSCAN algorithm is used to cluster the sample points representing merchant addresses. When calculating the distance matrix between sample points, the distance is defined as the spherical distance between two sample points (e.g. km); the algorithm parameters MinPts=2, and the initial value of ∈ is 0.5; for the DBSCAN clustering result, the constraint is applied: the merchant address sample points in the same class are within the same site, i.e. the merchant address sample points in the same class and within the same site can be merged into a merchant cluster.
[0118] Step 323: Cluster effect evaluation. Calculate the distance between each sample point in the merchant cluster and the average distance avg_dis, and measure the size of the merchant cluster; based on the business scenario, a threshold value (which can be configured and adjusted) is set, when the average distance is less than the threshold value, the merchant cluster can be used as an analyzed resident point, and the resident point and the included merchant address cluster are included in the analyzed resident point list, and the information point (Point of Information, POI) of the resident point can also be output in combination with the geographic location information. POI is any meaningful point on the map without geographical significance, such as stores, bars, gas stations, hospitals, stations, etc.; for the noise points obtained by clustering, each noise point is regarded as an analyzed resident point, and is also included in the analyzed resident point list.
[0119] Step 324: adjust parameters, re-cluster. For the sample points contained in the merchant cluster that does not pass the evaluation in step 323, reduce the value of by a step size, i.e., re-establish the DBSCAN model clustering, and repeat step 323 until or the clustering merchant cluster passes the evaluation. Finally output the list of analysis site points, for example as shown in Table 3.
[0120] Table 3: Example of analysis site point list
[0121]
[0122]
[0123] Step 330, time series feature mining. The target of mining includes one or more of the seasonal features of the site point shipment volume and the future shipment volume change features of the site point.
[0124] The seasonality of time series data is generally observed by observing the time series graph. If there is periodic fluctuation in the time series graph, there is seasonality, which can be further characterized by statistical calculation of the seasonality index. However, in the site business, the site point is large in magnitude, and the seasonal analysis method of a single time series of a single site point is not applicable, and the seasonal characteristics of a single site point are not the focus of mining. Therefore, this scheme starts from the perspective of shape-based time series clustering, and labels the seasonal characteristics of the site point according to the seasonal characteristics of the large category after clustering the site point according to the time series data. Since the seasonality of site point recommendation mainly evaluates whether there is a sudden increase / decrease in volume and its duration, this scheme realizes shape-based time series clustering through K-shape algorithm. The principle of K-shape algorithm is similar to that of K-means algorithm, but K-shape improves the distance calculation method, i.e. defines shape-based distance (SBD) through cross-correlation method, and optimizes the method of calculating the centroid, supports time series amplitude scaling and translation invariance, and has high computing efficiency.
[0125] In Figure 5 : (a) represents a time series data A, (b) represents another time series data B, (c) represents the similarity of time series data A and time series data B calculated according to K-means, (d) represents the similarity of time series data A and time series data B calculated according to K-shape algorithm, (e) represents an example of K-shape algorithm clustering result, cluster 1 (Cluster 1), cluster 2 (Cluster 2), cluster 3 (Cluster 3).
[0126] For time series data prediction, considering the difficulty of parameter tuning and prediction effect, Prophet algorithm is selected. Prophet algorithm model is based on time series decomposition, which consists of four parts: growth trend term g(t), periodic term s(t), special holiday effect term h(t) and error term ∈ t There are two basic forms:
[0127] Additive model: y = g(t) + s(t) + h(t) + ∈ t
[0128] Multiplicative model: y = g(t) * s(t) * h(t) * ∈ t
[0129] Among them, the multiplicative model can be converted to the additive model by logarithmic transformation, that is:
[0130] lny(t) = lng(t) + lns(t) + lnh(t) + ln∈ t
[0131] The trend function g(t) is constructed by piecewise logistic regression and piecewise linear, respectively, and two trend terms. The specific forms are as follows:
[0132] Saturation trend term:
[0133] Linear trend term: g(t) = (k + a(t)δ) * t + (m + a(t)γ)
[0134] Where C(t) is a function of time t, representing the upper limit capacity that the saturation trend term can reach, i.e. the maximum value. k + a(t) T δ represents the growth rate at time t, when k + a(t) T δ is greater than 0, it means upward trend, otherwise downward, the absolute value is larger, the faster the speed of this trend; k is a constant obeying normal distribution, representing the basic growth rate, δ = (δ1, δ2, …, δ s ) T , δ j corresponds to the growth rate change at time j, a(t) = (a1(t), a2(t), …, a s (t)), where s j is the time corresponding to the jth moment. m + a(t)γ is the translation term, where m is a constant obeying normal distribution, γ = (γ1, γ2, …, γ s ) T , in the saturation trend term γ j = -s j δ j.
[0135] The seasonal function s(t) is based on a Fourier series and is modeled as follows:
[0136]
[0137] where P represents the period (year: 365.25, week: 7), N represents the number of approximations used, i.e., N in the Fourier expansion, and the larger N is, the more refined the approximation is. β = (a1, b1, a2, b2, …, aN, bN) are the Fourier coefficients. n n T , Normal represents a normal distribution.
[0138] The model form of the holiday effect function h(t) is as follows:
[0139]
[0140] where L represents the number of holidays, D i represents the date set of such holidays in the past and future, is the one-hot-encoding of this holiday, κ = (κ1, κ2, …, κL) represents the impact of the corresponding holiday. L T , The larger v is, the greater the impact of the holiday on the model is, and vice versa. κ i represents the impact range of the corresponding holiday.
[0141] The error term is subject to a normal distribution and is represented as
[0142] In Figure 6 : (a) represents a schematic diagram of the prediction result of the Prophet algorithm, and (b) represents a schematic diagram of the effect result of each decomposition term of the Prophet algorithm.
[0143] Since the shipment volume of the to-be-analyzed resident point may have a long tail effect, and in order to optimize the clustering and prediction efficiency, the to-be-analyzed resident point needs to be screened before time series feature mining.
[0144] The time series feature mining process is described below.
[0145] Step 331: Taking “merchant code-sending address attribution site ID-sending address longitude-sending address latitude” as the primary key, the merchant shipment of the to-be-analyzed resident point is associated, and the shipment of each to-be-analyzed resident point in the specified historical period (such as the past two years) is summarized in day / month dimensions, respectively, to construct the day / month dimension shipment time series data of the corresponding merchant cluster. Thus, the to-be-analyzed resident point shipment time series data is obtained.
[0146] Step 332: Based on the time series data of the to-be-analyzed resident point order volume, filter out long-tail resident points that far fail to meet the resident standard, and obtain the to-be-decided resident point and its order volume time series data.
[0147] Step 333: Based on the to-be-decided resident point order volume time series data, time series clustering and time series prediction can be realized in parallel, of course, one of them can also be realized.
[0148] Time series clustering (333a): First, calculate the monthly order volume proportion of each to-be-decided resident point in the current year (unit: %), obtain the modified monthly dimension order volume time series data of each to-be-decided resident point, input K-shape model to cluster each to-be-decided resident point, and obtain to-be-decided resident point clusters belonging to different time series pattern categories.
[0149] Time series prediction (333b): Based on the daily dimension order volume time series data of the to-be-decided resident point, establish a Prophet algorithm to predict the daily dimension order volume of the future 3 months for each to-be-decided resident point. The seasonality factor of the Prophet algorithm model considers annual seasonality, weekly seasonality, and set holidays, such as but not limited to: New Year's Day, Spring Festival, promotion day, etc.
[0150] Step 334: Time series evaluation.
[0151] For the time series clustering result, the average number of delivery months per year, the average monthly order volume proportion in the current year, the peak value number, and other seasonal characteristics that need to be considered for resident point recommendation are calculated for each to-be-decided resident point cluster, and are assigned to the to-be-decided resident points in the to-be-decided resident point cluster.
[0152] For the time series prediction result, for example, predict: ① the monthly dimension order volume characteristics of the to-be-decided resident point in the future specified period (such as the next three months), ② the daily dimension delivery frequency characteristics of the to-be-decided resident point, such as x days / month, ③ the future change trend characteristics of the to-be-decided resident point order volume: based on the Prophet prediction trend item data, do least squares linear fitting, and extract the linear trend slope.
[0153] Based on the to-be-decided resident point time series characteristic information table finally output by the above steps, an example is shown in Table 4.
[0154] Table 4: Example of to-be-decided resident point time series characteristic information
[0155]
[0156]
[0157] Step 340, resident point recommendation.
[0158] The step takes the existing field point operation as one of the reference standards for the field point to be decided, and combines the seasonal characteristics and future time characteristics of the field point to be decided to recommend the field point.
[0159] The embodiment proposes a scheme of mining reference standards based on the operation of the existing field point and applying the reference standards, which is realized by a collaborative filtering method. The collaborative filtering is a recommendation algorithm based on the idea that "similar objects have similar properties". The algorithm calculates the similarity between objects and the similarity between users to predict the scores or preferences of users for unknown objects, so as to recommend the objects of interest of the users to the users. The commonly used similarity measures include the following.
[0160] (1) Euclidean distance. The Euclidean distance between two points P(p1, p2,…, pn) and Q(q1, q2,…, qn) is:
[0161]
[0162] The Euclidean distance is non-negative, and needs to be converted twice to make the value between -1 and 1 or between 0 and 1. The Euclidean distance is not suitable for Boolean vectors.
[0163] (2) Cosine similarity / adjusted cosine similarity.
[0164] The cosine similarity between two points P(p1, p2,…, pn) and Q(q1, q2,…, qn) is:
[0165]
[0166] The cosine similarity is between -1 and 1, and is independent of the length of the vector and not sensitive to the absolute value. In practical applications, the adjusted cosine similarity can be calculated.
[0167] The calculation process of the adjusted cosine similarity includes: first, calculating the mean value of each dimension of the vector, then subtracting the mean value of each dimension of each vector, and finally calculating the cosine similarity based on the adjusted vector. For example, assuming that there are two sample points P(p1, p2,…, pn) and Q(q1, q2,…, qn), the adjusted cosine similarity of the two sample points P and Q is: wherein Other sample points can be expanded in the same way.
[0168] (3) Pearson correlation coefficient. The Pearson correlation coefficient between two points P(p1, p2,…, pn) and Q(q1, q2,…, qn) is:
[0169]
[0170] The Pearson correlation coefficient takes a value between -1 and 1, and since the coefficient measures whether the trend of the two point vectors is consistent, it is not suitable for Boolean vectors.
[0171] When the present solution builds a decision recommendation model for a resident point based on a collaborative filtering algorithm, the actual operation mode of the resident point recommendation business is combined, and the user considered when making a decision recommendation is abstracted as 1. The similarity between the to-be-decided resident point and the existing resident point (the resident point currently operated) is used to make a decision recommendation. Since the features include real values and Boolean values, and in order to consider the influence of the absolute value size, the similarity of the present embodiment selects the adjusted cosine similarity measure.
[0172] The process of resident point recommendation is described below.
[0173] Step 341: Combine the pickup merchant order detail table to filter the existing resident points with stable operation. Thus, the existing resident points with large and stable order quantities are filtered out.
[0174] Step 342: Feature construction. For the filtered existing resident points and the to-be-decided resident points, features that can be used to evaluate the similarity between the two (these features can affect the decision of the resident point) are generated, and the features are filtered through a tree model to retain important features (these important features have a greater impact on the decision of the resident point) for constructing the feature vector required for calculating the cosine similarity.
[0175] The features used to calculate the similarity include numerical features and categorical features, etc. The numerical features include time series features and non-time series features.
[0176] Among the numerical features, for the feature vectors of the time series features (such as monthly order quantity, delivery frequency, etc.), the existing resident points can be filled based on actual operation data statistics, and the to-be-decided resident points can be filled with mined time series feature data. For the feature vectors of non-time series features (such as signing discount, cargo volume, and labor efficiency, etc.), the existing resident points and the to-be-decided resident points can be filled based on actual operation data statistics.
[0177] Among the categorical features, such as merchant categories and express product types, it is necessary to split and convert them into feature vectors with logical values, such as merchant category 1, merchant category 2, …, and merchant category n.
[0178] Step 343: Collaborative filtering. Based on the constructed feature vector, the adjusted cosine similarity value is calculated for each station to be decided and each existing station in stable operation. The proportion of the waybill volume of each existing station in stable operation to the total number of waybill volume of all station in stable operation is then counted. Finally, this proportion is used as the weight to calculate the weighted adjusted cosine similarity of each station to be decided. This weighted similarity is used to measure the overall similarity between the station to be decided and all existing station in stable operation. A threshold is set based on business considerations (this is configurable and adjustable). Stations to be decided whose weighted adjusted cosine similarity meets the threshold can proceed to the next decision.
[0179] Step 344: Comprehensive recommendation. Comprehensively consider the temporal characteristics of the stationed points to be decided, conduct a secondary evaluation on the stationed points to be decided outputted in step 343, set a threshold based on business considerations (the threshold is configurable and adjustable), filter out the stationed points with a sudden drop in future order volume or stationed points with obvious seasonality, and the remaining stationed points to be decided can be recommended as stationed points. Mark the key feature information of the recommended stationed points, such as the affiliated site, merchant list, future order volume, delivery frequency, estimated cost savings, etc. Finally, output the recommended stationed point information table. The recommended stationed point information table can be pushed to the stationed point management system to drive the establishment of stationed points.
[0180] The on-site point mining and recommendation scheme is implemented based on clustering and prediction algorithms, which has at least the following effects.
[0181] (1) In terms of space, geographic information data is introduced, and a density-based clustering algorithm such as DBSCAN is used. A clustering result loop iteration link is set up to fully explore the on-site points to be analyzed and the corresponding merchant list, and obtain the on-site points with low difficulty in collection operations and whose jurisdiction can be regarded as the "same shipping address". A on-site point represents a virtual location point / range suitable for switching to the on-site point and the corresponding merchant list. This makes the identification process of the on-site merchants that should be switched more accurate and reasonable, and improves the coverage of recommended merchants.
[0182] (2) In terms of time, the time series clustering / forecasting algorithm is used to construct the seasonal characteristics and future time series characteristics of the waybill volume of the stationed points based on the time series clustering / forecasting results. The stationed points are recommended based on the comprehensive consideration of the historical waybill volume and future changes in the waybill volume of the stationed points, and the recommendation strategy is optimized to reduce the probability of problems such as employee redundancy and increased management difficulty after the establishment of the stationed points.
[0183] (3) In terms of decision-making strategy, we expand the consideration of multiple factors affecting on-site switching, and establish a recommendation decision-making strategy based on the existing stable operating on-site points as the standard through the collaborative filtering algorithm similarity measurement, so that the recommendation results are more reasonable and in line with the actual operating conditions, thereby improving the accuracy of on-site point identification.
[0184] (4) On the system, the recommendation process automation is realized, and the recommendation standard iteration is simplified. Specifically, through four steps of basic data, merchant address clustering, time sequence feature mining, and on-site point recommendation, the on-site point automatic mining and recommendation are realized; in terms of the on-site point recommendation standard, the simple adjustment of the merchant address clustering threshold and the on-site point recommendation threshold can realize iteration, so that the iteration is simplified.
[0185] Figure 7 A structural schematic diagram of an on-site point mining apparatus of some embodiments of the present disclosure is shown. As shown in the figure, the on-site point mining apparatus 700 of the embodiment includes a memory 710 and a processor 720 coupled to the memory 710, and the processor 720 is configured to execute the on-site point mining method in any of the foregoing some embodiments based on instructions stored in the memory 710. Figure 7
[0186] The memory 710 may, for example, include a system memory, a fixed nonvolatile storage medium, etc. The system memory, for example, stores an operating system, an application program, a boot loader, and other programs, etc.
[0187] The processor 720 may, for example, be implemented in the form of a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete hardware component, or a combination thereof.
[0188] The on-site point mining apparatus 700 may, for example, further include an input / output interface 730, a network interface 740, a storage interface 750, etc. These interfaces 730, 740, 750 and the memory 710 and the processor 720 may, for example, be connected through a bus 760. The input / output interface 730 provides a connection interface for display, mouse, keyboard, touch screen, and other input / output devices. The network interface 740 provides a connection interface for various networking devices. The storage interface 750 provides a connection interface for external storage devices such as SD cards and U disks. The bus 760 may, for example, use any of a variety of bus structures. For example, the bus structure includes but is not limited to an industry standard architecture (ISA) bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus.
[0189] Figure 8 A structural schematic diagram of a field point mining device is shown to illustrate some embodiments of the present disclosure. As shown, the field point mining device 800 of this embodiment includes a module for performing a field point mining method. Figure 8
[0190] The clustering module 810 is configured to obtain the geographic locations of the shipment addresses of each non-stationed customer; construct a sample point set with each geographic location of each shipment address of each non-stationed customer as a sample point; cluster each sample point in the sample point set to obtain each sample point cluster; and determine a candidate field point based on each sample point cluster.
[0191] The clustering module 810 is configured to cluster each sample point in the sample point set using a density-based clustering algorithm, and impose a constraint on the clustering result, so that sample points of the same type and in the same station are aggregated into a sample point cluster.
[0192] The clustering module 810 is configured to calculate a distance mean based on the distances between sample points in each sample point cluster; and determine a sample point cluster with a distance mean less than a distance threshold as a candidate field point.
[0193] The mining module 820 is configured to determine a field point based on the candidate field point.
[0194] The first mining unit 821 is configured to determine a field point based on the number of shipments of the candidate field point.
[0195] The second mining unit 822 is configured to remove a candidate field point with a long tail feature based on the time series data of the number of shipments of the candidate field point; and determine a field point based on the remaining candidate field points.
[0196] The third mining unit 823 is configured to determine a field point based on one or more of the seasonal characteristics and future time series characteristics of the number of shipments of the candidate field point.
[0197] The third mining unit 823 is configured to remove a candidate field point with seasonal fluctuations based on the seasonal characteristics of the number of shipments of the candidate field point, and / or remove a candidate field point with a future number of shipments below a threshold based on the future time series characteristics of the number of shipments of the candidate field point; and determine a field point based on the remaining candidate field points.
[0198] The third mining unit 823 is configured to determine the seasonal characteristics of the number of shipments of the candidate field point, including: obtaining each field point cluster with different time series pattern categories based on the time series data of the number of shipments of each candidate field point using a shape-based time series clustering algorithm; and determining the seasonal characteristics of each field point cluster as the seasonal characteristics of the number of shipments of the candidate field points in the each field point cluster.
[0199] The third mining unit 823 is configured to determine the future time sequence feature of the shipment volume of the candidate resident point, including: establishing a time sequence prediction algorithm, such as a Prophet algorithm, for each candidate resident point according to the time sequence data of the shipment volume of each candidate resident point, to predict the future time sequence feature of the shipment volume of each candidate resident point, wherein the time sequence prediction algorithm, such as the Prophet algorithm, includes a growth trend item, a periodic item, a holiday effect, and an error item.
[0200] The fourth mining unit 824 is configured to determine the resident point according to the similarity between the candidate resident point and the existing resident point set.
[0201] The fourth mining unit 824 is configured to calculate the similarity between each candidate resident point and each existing resident point in the existing resident point set, determine the weight of each existing resident point as the proportion of the shipment volume of the existing resident point to the total shipment volume of the existing resident point set, and perform weighted calculation according to the similarity between each candidate resident point and each existing resident point in the existing resident point set and the weight of each existing resident point to obtain a weighted similarity as the similarity between each candidate resident point and the existing resident point set.
[0202] The fourth mining unit 824 is configured to calculate the similarity between the feature vector of each candidate resident point and the feature vector of each existing resident point in the existing resident point set as the similarity between each candidate resident point and each existing resident point in the existing resident point set, wherein the plurality of features are determined according to the plurality of features affecting the resident point decision.
[0203] The fourth mining unit 824 is configured to filter the plurality of initial features affecting the resident point decision through a tree model, and construct a feature vector of the plurality of features based on the plurality of features with a required importance degree.
[0204] Some embodiments of the present disclosure provide a computer readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the resident point mining method.
[0205] Some embodiments of the present disclosure provide a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the resident point mining method.
[0206] It should be noted that in the technical solutions of the present disclosure, the collection, collection, updating, analysis, processing, use, transmission, storage, etc. of user personal information involved in the technical solutions comply with relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data, and the security of user personal information, network security and national security are maintained.
[0207] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more non-transitory computer-readable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer program code.
[0208] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0209] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0210] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0211] The above merely provides the preferred embodiments of the present disclosure, and is not intended to limit the present disclosure. Any modification, equivalent replacement, improvement, and the like made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for determining a field point, comprising: obtaining geographical locations of shipment addresses of each non-fielded customer; constructing a sample point set by taking the geographical location of each shipment address of each non-fielded customer as a sample point; clustering each sample point in the sample point set to obtain each sample point cluster; determining a candidate field point according to each sample point cluster; and determining a field point according to the candidate field point.
2. The method of claim 1, wherein, The clustering each sample point in the sample point set to obtain each sample point cluster comprises: clustering each sample point in the sample point set by using a density-based clustering algorithm, and imposing a constraint on the clustering result, so that sample points of the same type and in the same site are aggregated into a sample point cluster.
3. The method of claim 1, wherein, The determining a candidate field point according to each sample point cluster comprises: calculating a distance mean based on distances between sample points in each sample point cluster; and determining a sample point cluster with a distance mean less than a distance threshold as a candidate field point. 4.The method of any one of claims 1-3, wherein the determining a field point according to the candidate field point comprises: determining a field point according to a shipment volume of the candidate field point.
5. The method of claim 4, wherein, The determining a field point according to the shipment volume of the candidate field point comprises: removing a candidate field point with a long tail feature according to time series data of the shipment volume of the candidate field point; and determining a field point according to the remaining candidate field point. 6.The method of any one of claims 1-5, wherein the determining a field point according to the candidate field point comprises: removing a candidate field point with seasonal fluctuations according to a seasonal feature of the shipment volume of the candidate field point, and / or removing a candidate field point with a future shipment volume less than a threshold according to a future time series feature of the shipment volume of the candidate field point; and determining a field point according to the remaining candidate field point.
7. The method of claim 6, further comprising: The determining a seasonal feature of the shipment volume of the candidate field point comprises: obtaining each field point cluster with different time series pattern categories by using a shape-based time series clustering algorithm according to time series data of the shipment volume of each candidate field point; and determining a seasonal feature of each field point cluster as a seasonal feature of the shipment volume of a candidate field point in the each field point cluster. Alternatively, the seasonal feature of each field point cluster comprises one or more of an average number of shipment months per year, an average proportion of monthly shipment volume in annual shipment volume, and a number of peaks.
8. The method of claim 6, further comprising: The determining a future time series feature of the shipment volume of the candidate field point comprises: establishing a time series prediction algorithm for each candidate field point according to time series data of the shipment volume of the each candidate field point, to predict a future time series feature of the shipment volume of the each candidate field point, wherein the time series prediction algorithm comprises a growth trend term, a periodic term, a holiday effect term, and an error term. Alternatively, the future time series feature of the shipment volume of the candidate field point comprises one or more of a future shipment volume of the candidate field point, a shipment frequency feature of the candidate field point, and a future variation trend feature of the shipment volume of the candidate field point. 9.The method of any one of claims 1-8, wherein the determining a field point according to the candidate field point comprises: determining a similarity between the candidate field point and an existing field point set. Determine the station point according to the similarity between the candidate station point and the existing station point set.
10. The method of claim 9, wherein, Determine the similarity between the candidate station point and the existing station point set, comprising: Calculate the similarity between each candidate station point and each existing station point in the existing station point set; Determine the weight of the existing station point as the proportion of the number of waybills of the existing station point to the total number of waybills of the existing station point set; According to the similarity between each candidate station point and each existing station point in the existing station point set and the weight of each existing station point, weighted calculation is performed to obtain the weighted similarity as the similarity between each candidate station point and the existing station point set.
11. The method of claim 10, wherein, The calculation of the similarity between each candidate station point and each existing station point in the existing station point set comprises: Calculate the similarity between the feature vector of each candidate station point and the feature vector of each existing station point in the existing station point set, as the similarity between each candidate station point and each existing station point in the existing station point set, wherein the plurality of features are determined according to the plurality of features affecting the station point decision.
12. The method of claim 11, wherein, The plurality of features are obtained by screening the plurality of initial features affecting the station point decision by using a tree model. Alternatively, the plurality of features include one or more of the time series feature of the number of waybills, the signing discount feature, the cargo volume feature, the user category feature, the express product type feature.
13. A field point mining apparatus comprising: Memory; And a processor coupled to the memory, the processor is configured to execute the station point mining method according to any one of claims 1-12 based on the instructions stored in the memory.
14. A field point mining apparatus comprising: The module for executing the station point mining method according to any one of claims 1-12.
15. A computer readable storage medium storing a computer program, which is executed by a processor to implement the steps of the station point mining method according to any one of claims 1-12.
16. A computer program product comprising a computer program, which is executed by a processor to implement the steps of the station point mining method according to any one of claims 1-12.