Logistics site selection method and system based on combination of density peak clustering and Gaussian model

By combining density peak clustering and Gaussian models, the system automatically analyzes multi-dimensional data in logistics site selection, optimizes the location of logistics centers, and solves the problems of insufficient consideration of multi-dimensional factors and poor dynamic adaptability in traditional methods, thus achieving more scientific and reasonable logistics site selection decisions.

CN120996485APending Publication Date: 2025-11-21SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511154760.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional logistics site selection methods struggle to fully consider multi-dimensional factors, lack data support, and have poor dynamic adaptability, leading to unscientific and unreasonable site selection decisions.

Method used

A method combining density peak clustering and Gaussian model is adopted. By calculating local density through information entropy and Gaussian mixture model, multi-dimensional data in logistics site selection problem is automatically analyzed, objective function is constructed and constraints are set to optimize site selection scheme.

Benefits of technology

It improves the scientific and rational nature of logistics site selection, enabling it to adapt to dynamic changes and optimize site selection schemes to reduce operating costs and improve efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996485A_ABST
    Figure CN120996485A_ABST
Patent Text Reader

Abstract

The invention provides a logistics site selection method and system based on combination of density peak clustering and a Gaussian model, and belongs to the field of logistics planning and data analysis. The method comprises the steps of collecting multi-dimensional data in logistics site selection, and performing preprocessing; carrying out clustering analysis on the preprocessed multi-dimensional data by adopting an information entropy-based density peak value clustering algorithm and a Gaussian mixture model combined clustering algorithm, and extracting a clustering center of the data as a potential logistics site selection point; and according to a preset logistics site selection target, constructing a target function and setting a constraint condition, and solving an optimal logistics site selection point from the potential logistics site selection points through an optimization algorithm. Therefore, the internal structure and rule behind the data are disclosed. The limitation that a traditional method depends on a single factor is avoided, the objectivity and reliability of the site selection scheme are improved, the site selection scheme can be rapidly adjusted through data updating and re-clustering, and the optimality is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of logistics planning and data analysis, and in particular relates to a logistics site selection method and system based on the combination of density peak clustering and Gaussian model. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] In the logistics industry, the location of a logistics center is a crucial decision-making process, directly impacting the layout of the logistics network, operational efficiency, cost control, and customer service quality. Simply put, logistics site selection involves choosing one or more suitable locations in geography to establish a logistics center or distribution center based on the needs of logistics operations, enabling the efficient collection, storage, and distribution of goods. Appropriate logistics site selection can save operating costs, and optimized site selection can reduce inventory backlog and lower warehousing costs. It also helps improve logistics efficiency, enabling faster response to customer needs and providing more timely and accurate delivery services, thereby enhancing customer service quality.

[0004] Traditional logistics site selection methods often rely on human experience and simple geographical analysis, such as considering the convenience of the location and transportation conditions. However, these methods have the following problems: (1) Insufficient consideration of multi-dimensional factors: Traditional methods often fail to fully consider the complexity and dynamism of multi-dimensional factors (such as geographical location, traffic conditions, customer needs, land costs, etc.), resulting in unscientific and unreasonable site selection decisions.

[0005] (2) Lack of data support: Traditional methods often lack sufficient data support, making it difficult to quantitatively evaluate and optimize site selection schemes.

[0006] (3) Poor dynamic adaptability: The logistics industry environment is constantly changing, and factors such as customer needs and traffic conditions are also constantly changing. Traditional methods are difficult to adapt to these dynamic changes, making it difficult to continuously optimize the site selection scheme. Summary of the Invention

[0007] To overcome the shortcomings of the prior art, this invention provides a logistics location selection method and system based on the combination of density peak clustering and Gaussian model. By using cluster analysis technology to automatically analyze multi-dimensional data in the logistics location selection problem, and combined with specific objectives (such as minimizing costs, maximizing coverage, etc.), the method optimizes the location scheme of logistics centers, improves logistics efficiency, and reduces operating costs.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a logistics location selection method based on a combination of density peak clustering and Gaussian model; Logistics location selection methods based on the combination of density peak clustering and Gaussian models include: Collect multi-dimensional data on logistics site selection and preprocess it; A clustering algorithm combining density peak clustering based on information entropy and Gaussian mixture model is used to perform cluster analysis on preprocessed multidimensional data, and the cluster centers of the data are extracted as potential logistics site selection points. Based on the preset logistics location selection objectives, an objective function is constructed and constraints are set. The optimal logistics location is then determined from potential logistics location points using an optimization algorithm.

[0009] As a further technical solution, the multi-dimensional data includes geographical location, traffic conditions, customer needs, and land costs.

[0010] As a further technical solution, the preprocessing process includes data cleaning, data normalization, and data dimensionality reduction.

[0011] As a further technical solution, a clustering algorithm combining density peak clustering based on information entropy and Gaussian mixture model is used to perform cluster analysis on the preprocessed multi-dimensional data, extracting the cluster centers as potential logistics location points, including: Calculate the local density of each data point based on the Gaussian kernel and information entropy; For each data point, find the point with the smallest distance among points with higher density, and calculate the distance between the high-density points; A decision graph is plotted using the distance between local density points and high density points as the x-axis and y-axis, respectively, and the decision graph is used to identify cluster centers. Points that were not selected as cluster centers are assigned to the cluster containing the nearest and densest point. The cluster centers and covariance matrix of the Gaussian mixture model are initialized using the cluster centers and remaining points obtained by the density peak clustering algorithm based on information entropy. The cluster centers of the extracted data are then used as potential logistics location points through iterative optimization.

[0012] As a further technical solution, the preset logistics site selection objectives include minimizing total cost, maximizing coverage, minimizing maximum delivery distance, and balancing cost and coverage.

[0013] As a further technical solution, the constraints include demand point allocation constraints, site selection point service capacity constraints, delivery distance limits, site selection point quantity limits, site selection point and demand point association constraints, land cost limits, coverage constraints, and site selection point capacity constraints.

[0014] As a further technical solution, the optimization algorithm includes at least one of integer linear programming, genetic algorithm, simulated annealing algorithm or particle swarm optimization algorithm.

[0015] The second aspect of this invention provides a logistics location system based on a combination of density peak clustering and Gaussian model.

[0016] A logistics location selection system based on a combination of density peak clustering and Gaussian model includes: The data acquisition module is configured to collect multi-dimensional data in logistics site selection and perform preprocessing. The clustering analysis module is configured to use a clustering algorithm that combines density peak clustering based on information entropy with Gaussian mixture model to perform clustering analysis on preprocessed multi-dimensional data and extract the cluster centers of the data as potential logistics location points. The location optimization module is configured to: construct an objective function and set constraints based on the preset logistics location objectives, and solve for the optimal logistics location from potential logistics location points through an optimization algorithm. A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the logistics location method based on the combination of density peak clustering and Gaussian model as described in the first aspect of the present invention.

[0017] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the logistics location method based on the combination of density peak clustering and Gaussian model as described in the first aspect of the present invention.

[0018] The above one or more technical solutions have the following beneficial effects: This invention, by comprehensively considering multiple dimensions such as geographical location, traffic conditions, customer demand, and land costs, groups data points with similar characteristics into one category, thereby revealing the underlying structure and patterns of the data. This helps to select logistics center locations more scientifically and rationally. It employs a density peak clustering algorithm based on information entropy to automatically identify cluster centers, without requiring a pre-specified number of clusters, and can adaptively discover natural cluster structures in the data. A Gaussian mixture model is used to probabilistically model the clustering results, optimizing the cluster distribution and boundaries, and improving the representativeness of the selected sites.

[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0021] Figure 1 This is a flowchart of the method in the first embodiment.

[0022] Figure 2 This is a system structure diagram of the second embodiment. Detailed Implementation

[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0025] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0026] Example 1 This embodiment discloses a logistics location selection method based on the combination of density peak clustering and Gaussian model; like Figure 1 As shown, the logistics location selection method based on the combination of density peak clustering and Gaussian model includes: Step S1: Collect multi-dimensional data in logistics site selection and perform preprocessing; Step S2: A clustering algorithm combining density peak clustering algorithm based on information entropy and Gaussian mixture model is used to perform cluster analysis on the preprocessed multi-dimensional data, and the cluster centers of the data are extracted as potential logistics site selection points. Step S3: Based on the preset logistics location target, construct the objective function and set the constraints, and solve for the optimal logistics location point from the potential logistics location points through the optimization algorithm.

[0027] Specifically, it also includes the following: In step S1, multi-dimensional data on logistics site selection is collected, including geographical location, traffic conditions, customer demand, and land costs. Data sources may include Geographic Information Systems (GIS), traffic flow data, and market research reports.

[0028] Furthermore, the acquired multi-dimensional data is preprocessed, including data cleaning, data normalization, and data dimensionality reduction.

[0029] Data cleaning is used to remove duplicate, erroneous, or invalid data.

[0030] Data normalization uses Z-Score normalization, which transforms the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. The formula is expressed as:

[0031] in: These are the original data values; It is the mean of the dataset; It is the standard deviation of the dataset; These are the normalized data values.

[0032] Principal Component Analysis (PCA) can be used to reduce data dimensionality and computational complexity. PCA is a linear transformation method for reducing data dimensionality. It maps high-dimensional data to a low-dimensional space by finding the direction of maximum variance (principal component) in the data, while preserving the main information of the data.

[0033] In step S2, the local density is first obtained using a density peak clustering algorithm based on information entropy. Local density Used to measure data points The density of the surrounding environment, combined with information entropy, enhances interpretability.

[0034] The information entropy calculation process is as follows:

[0035] in, Data points Within its neighborhood, it belongs to the first The probability distribution of each feature.

[0036] Then, the local density is:

[0037] in: It is a point and points The distance between them; It is the cutoff distance; It is a point Information entropy is used to reflect the complexity of local feature distribution.

[0038] Furthermore, for each data point, find the point with the smallest distance among points with higher density, and calculate the distance between these high-density points. The distance between the high-density points is... Used to measure points Minimum distance to a point with a higher density:

[0039] Local density and distance to high-density points are used as the x and y axes, respectively. Decision graphs are drawn and used to identify cluster centers, which typically have high local density. Larger distance between high-density points .

[0040] Points not selected as cluster centers are assigned to the cluster containing the nearest, higher-density point. The cluster centers and covariance matrix of the Gaussian mixture model are initialized using the cluster centers obtained from the density peak clustering algorithm based on information entropy and the clustering results of the remaining points. The Gaussian mixture model assumes that the data is a mixture of multiple Gaussian distributions. For each cluster, the probability density function of the GMM is:

[0041] in: It is the first The mixing coefficients of a Gaussian distribution satisfy the following conditions: . It is the first The probability density function of a Gaussian distribution has a mean of . The covariance matrix is .

[0042] GMM uses the Expectation-Maximization (EM) algorithm to estimate parameters.

[0043] E-step (the expected step involves calculating whether each data point belongs to the first step) The posterior probability (responsibility value) of each cluster:

[0044] in, For a given data point Given the model parameters, data points Belongs to the The posterior probability of each cluster; For this is the first The mean of a Gaussian cluster; For the first The covariance matrix of a Gaussian variety; For this is the first The prior probability of a Gaussian variety is also called the mixing coefficient.

[0045] The M-step (maximization step) updates the model parameters based on the posterior probability:

[0046]

[0047]

[0048] Through iterative optimization, the cluster centers (mean of GMM) resulting from the combination of EDPC and GMM are used as potential logistics location points. These points have high local density and large distances between high-density points, while the probability distribution and covariance structure of the clusters are optimized through GMM.

[0049] In step S3, the preset logistics location selection objectives include minimizing total cost, maximizing coverage, minimizing maximum delivery distance, and balancing cost and coverage.

[0050] The objective function is constructed to minimize the total cost:

[0051] in, For binary variables, Indicates the selection of a site point ,otherwise . For binary variables, Indicate demand points From the site selection point Service, otherwise . For site selection Fixed costs (such as land costs, construction costs, etc.). For demand points To the site selection point The distance. For demand points To the site selection point The unit transportation cost (which may be related to distance). This is a set of demand points.

[0052] The objective function constructed to maximize coverage is:

[0053] in, For demand points Weights (such as demand, priority, etc.).

[0054] The objective function constructed to minimize the maximum delivery distance is:

[0055] The objective function constructed to balance cost and coverage is as follows:

[0056] in, These are weighting coefficients used to balance costs, transportation costs, and coverage.

[0057] Furthermore, constraints are set according to actual needs. These constraints include constraints on demand point allocation, site selection service capacity, delivery distance limits, site selection quantity limits, the relationship between site selection and demand points, land cost limits, coverage constraints, and site selection capacity constraints. The demand point allocation constraint stipulates that each demand point must be served by one and only one location point:

[0058] Site selection service capacity constraint: The service capacity of each site selection point is limited.

[0059] in, For demand points Demand; For site selection The maximum service capacity.

[0060] The delivery distance is limited to the maximum allowable distance between the demand point and the selected location. :

[0061] The number of site selection points is limited to a maximum of [number missing]. :

[0062] The association constraint between the location point and the demand point is that only the selected location point can serve the demand point:

[0063] Land cost is limited to a total land cost that cannot exceed the budget. :

[0064] The coverage constraint ensures that all or high-priority requirements are covered:

[0065] in, This is a set of high-priority requirements.

[0066] The site selection capacity constraint is that the capacity of each site selection point cannot exceed its maximum capacity. :

[0067] in, For demand points The volume of goods.

[0068] The objective function is solved using optimization algorithms (such as linear programming, integer programming, heuristic algorithms, etc.). Among them, integer linear programming (ILP) models the location problem as an integer linear programming problem and solves it using commercial solvers (such as CPLEX, Gurobi) or open-source solvers (such as GLPK).

[0069] Heuristic algorithms include genetic algorithms, simulated annealing algorithms, and particle swarm optimization algorithms. Genetic algorithms optimize site selection by simulating natural selection and genetic mechanisms. Simulated annealing algorithms avoid getting trapped in local optima by simulating the annealing process. Particle swarm optimization algorithms find the optimal solution by simulating the movement of particles in the solution space. In this embodiment, one of the above optimization algorithms can be selected according to requirements, and none of them constitutes a limitation on this solution. Furthermore, the above optimization algorithms are all conventional techniques in this field and will not be described in detail here.

[0070] Mixed integer programming (MIP) is used to linearize or use a mixed integer programming solver if the objective function or constraints contain nonlinear terms.

[0071] Furthermore, the site selection results are integrated into structured data formats (such as JSON and XML) for easier subsequent processing and storage. Visualization tools such as maps are used to display the results, providing a clear overview of the distribution of logistics centers. Based on the logistics site selection results, downstream tasks such as logistics network planning, cost-benefit analysis, and risk assessment can be performed.

[0072] Example 2 This embodiment discloses a logistics location selection system based on a combination of density peak clustering and Gaussian model; like Figure 2 As shown, the logistics location system based on the combination of density peak clustering and Gaussian model includes: The data acquisition module is configured to collect multi-dimensional data in logistics site selection and perform preprocessing. The clustering analysis module is configured to use a clustering algorithm that combines density peak clustering based on information entropy with Gaussian mixture model to perform clustering analysis on preprocessed multi-dimensional data and extract the cluster centers of the data as potential logistics location points. The location optimization module is configured to: construct an objective function and set constraints based on the preset logistics location objectives, and solve for the optimal logistics location from potential logistics location points through an optimization algorithm. Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0073] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the logistics location method based on the combination of density peak clustering and Gaussian model as described in Example 1.

[0074] Example 4 The purpose of this embodiment is to provide an electronic device.

[0075] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the logistics location selection method based on the combination of density peak clustering and Gaussian model as described in Example 1.

[0076] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0077] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0078] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A logistics location selection method based on a combination of density peak clustering and Gaussian model, characterized in that, include: Collect multi-dimensional data on logistics site selection and preprocess it; A clustering algorithm combining density peak clustering based on information entropy and Gaussian mixture model is used to perform cluster analysis on preprocessed multidimensional data, and the cluster centers of the data are extracted as potential logistics site selection points. Based on the preset logistics location selection objectives, an objective function is constructed and constraints are set. The optimal logistics location is then determined from potential logistics location points using an optimization algorithm.

2. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, The multi-dimensional data includes geographical location, traffic conditions, customer needs, and land costs.

3. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, The preprocessing process includes data cleaning, data normalization, and data dimensionality reduction.

4. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, A clustering algorithm combining density peak clustering based on information entropy and Gaussian mixture model is used to perform cluster analysis on preprocessed multi-dimensional data, extracting cluster centers as potential logistics location points, including: Calculate the local density of each data point based on the Gaussian kernel and information entropy; For each data point, find the point with the smallest distance among points with higher density, and calculate the distance between the high-density points; A decision graph is plotted using the distance between local density points and high density points as the x-axis and y-axis, respectively, and the decision graph is used to identify cluster centers. Points that were not selected as cluster centers are assigned to the cluster containing the nearest and densest point. The cluster centers and covariance matrix of the Gaussian mixture model are initialized using the cluster centers and remaining points obtained by the density peak clustering algorithm based on information entropy. The cluster centers of the extracted data are then used as potential logistics location points through iterative optimization.

5. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, The preset logistics site selection objectives include minimizing total cost, maximizing coverage, minimizing maximum delivery distance, and balancing cost and coverage.

6. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, The constraints include demand point allocation constraints, site selection service capacity constraints, delivery distance limits, site selection number limits, site selection and demand point association constraints, land cost limits, coverage constraints, and site selection capacity constraints.

7. The logistics location selection method based on the combination of density peak clustering and Gaussian model as described in claim 1, characterized in that, The optimization algorithm includes at least one of integer linear programming, genetic algorithm, simulated annealing algorithm, or particle swarm optimization algorithm.

8. A logistics location selection system based on a combination of density peak clustering and Gaussian model, characterized in that: include: The data acquisition module is configured to collect multi-dimensional data in logistics site selection and perform preprocessing. The clustering analysis module is configured to use a clustering algorithm that combines density peak clustering based on information entropy with Gaussian mixture model to perform clustering analysis on preprocessed multi-dimensional data and extract the cluster centers of the data as potential logistics location points. The location optimization module is configured to: construct an objective function and set constraints based on the preset logistics location objectives, and solve for the optimal logistics location from potential logistics location points through an optimization algorithm.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the logistics location method based on the combination of density peak clustering and Gaussian model as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the logistics location selection method based on the combination of density peak clustering and Gaussian model as described in any one of claims 1-7.

Citation Information

Cited By

  • Agricultural product circulation project site selection decision-making method and device and electronic equipment

    CN122243567A