A Machine Learning-Based Auxiliary Site Selection Method and Device

Through the auxiliary site selection method based on machine learning, an effective facility evaluation index system is built and an auxiliary site selection model is trained, which solves the problem of low efficiency in assisted site selection decision-making in the existing technology, and realizes a more efficient site selection decision-making process.

CN117609405BActive Publication Date: 2025-05-30INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311572701.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-05-30
Estimated Expiration
2043-11-22

AI Technical Summary

Technical Problem

The existing auxiliary site selection algorithms have a long time-consuming process due to the large variety of data, large scale and many processing processes, and the site selection decision-making efficiency is low, and cannot meet the needs of rapid development of the industry.

Method used

Using an auxiliary site selection method based on machine learning, we use various effective facilities in the area to be selected, build an evaluation index system, perform data normalization processing, create vector maps and grid processing, obtain sample address data sets, and train auxiliary site selection models for decision-making.

Benefits of technology

By optimizing the auxiliary site selection decision process, mining the patterns and relationships between data, the efficiency of site selection decisions is significantly improved, and the problem of long-term real-time decision-making process is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117609405B_ABST
    Figure CN117609405B_ABST
Patent Text Reader

Abstract

The present invention discloses an auxiliary site selection method and device based on machine learning, belonging to the technical field of data processing, and is used to solve the technical problems that the current auxiliary site selection algorithms involve a large variety of data, large scale, and many processing processes, resulting in a long time-consuming real-time decision-making process for auxiliary site selection and low site selection decision-making efficiency. The method includes: determining various types of effective facilities within the area to be selected, and constructing an evaluation index system for the effective facilities; normalizing the evaluation data of the effective facilities to obtain a comprehensive evaluation data set; creating a vector map of the area to be selected, and performing grid processing on the vector map of the area to be selected to obtain a grid distribution map; obtaining a sampling address data set in the grid distribution map through an address uniform sampling mechanism; determining a sampling address annotation data set based on the comprehensive evaluation data set and the sampling address data set; and determining a final auxiliary site selection model according to the sampling address annotation data set to make an auxiliary site selection decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to an auxiliary site selection method and device based on machine learning. Background Art

[0002] In industries such as construction projects, store planning, and park planning, site selection is of utmost importance. A suitable and effective site selection can bring great economic benefits to the industry. In the early stage of site selection decision-making, not only various relevant data need to be considered, but also the non-linear coupling relationships that are generally present and mutually related and influential among the relevant data. Therefore, site selection decisions usually require multiple data processing processes to process the data related to site selection decisions and the complex relationships among these data, and determine the site selection decision results through complex calculation processes.

[0003] Based on this, some auxiliary site selection algorithms have emerged to help people process complex data calculations. However, the currently available auxiliary site selection algorithms involve a large variety of data types, large scales, and multiple processing processes, resulting in a long time-consuming real-time decision-making process for auxiliary site selection, low site selection decision-making efficiency, and inability to meet the needs of the rapidly developing industry. Summary of the Invention

[0004] Embodiments of the present invention provide an auxiliary site selection method and device based on machine learning to solve the following technical problems: The currently available auxiliary site selection algorithms involve a large variety of data types, large scales, and multiple processing processes, resulting in a long time-consuming real-time decision-making process for auxiliary site selection and low site selection decision-making efficiency.

[0005] Embodiments of the present invention adopt the following technical solutions:

[0006] On the one hand, embodiments of the present invention provide an auxiliary site selection method based on machine learning, the method comprising: determining various types of effective facilities within the area to be site-selected based on the site selection requirements, and constructing an evaluation index system for the effective facilities;

[0007] Normalizing the evaluation data of the same type of effective facilities to obtain a comprehensive evaluation data set corresponding to each type of effective facility;

[0008] Obtaining the boundary longitude and latitude data set of the area to be site-selected, creating a vector map of the area to be site-selected, and performing grid processing on the vector map of the area to be site-selected to obtain a grid distribution map;

[0009] Obtaining a sampling address data set in the grid distribution map through an address uniform sampling mechanism;

[0010] Determining a sampling address annotation data set based on the comprehensive evaluation data set and the sampling address data set;

[0011] According to the distribution law of the point location labels of the sampling address points in the sampling address annotation dataset, select the corresponding auxiliary site selection model;

[0012] Through the sampling address annotation dataset, train the auxiliary site selection model and optimize it to obtain the final auxiliary site selection model, so as to make auxiliary site selection decisions through the final auxiliary site selection model.

[0013] In a feasible implementation manner, based on the site selection requirements, determine various effective facilities within the area to be site-selected, and construct an evaluation index system for effective facilities, specifically including:

[0014] Based on the site selection requirements and the usage nature of the facility to be site-selected, determine the effective facilities of several usage types that affect the site selection decision; among them, the effective facilities of the several usage types at least include: residential providing facilities, medical service facilities, educational service facilities, shopping service facilities, and transportation service facilities;

[0015] According to the relationship between the effective facilities and the facility to be site-selected, determine the type of influence of the effective facilities on the site selection decision; among them, the relationship between the effective facilities and the facility to be site-selected includes a dependence relationship and a competition relationship; the type of influence includes a positive influence and a negative influence;

[0016] According to the type of influence, determine the category of the role of the effective facilities; among them, the category of the role of the effective facilities corresponding to the positive influence is the positive effective facility, and the category of the role of the effective facilities corresponding to the negative influence is the negative effective facility;

[0017] Based on the facility attribute characteristics of the effective facilities of each usage type, construct an evaluation index system corresponding to the effective facilities of each usage type; among them, the evaluation indexes in the evaluation index system include: relevant indexes of the population living inside the facility, relevant indexes of the population receiving services inside the facility, relevant indexes of the service content of the facility, and relevant indexes of the internal attributes of the facility.

[0018] In a feasible implementation manner, normalize the evaluation data of the same type of effective facilities to obtain a comprehensive evaluation dataset corresponding to each type of effective facility, specifically including:

[0019] Obtain the spatial location data and evaluation data of various effective facilities within the area to be site-selected; among them, the spatial location data at least includes the longitude and latitude of the effective facilities; the evaluation data at least includes the evaluation index data of the effective facilities in their evaluation index systems;

[0020] Divide the evaluation indexes in the evaluation index system into positive evaluation indexes and negative evaluation indexes;

[0021] Based on the linear transformation algorithm, perform linear transformation on the evaluation data corresponding to the positive evaluation indicators and the evaluation data corresponding to the negative evaluation indicators respectively, and convert the evaluation data into standard evaluation data within the interval [0, 1] with the same direction of action, so as to obtain a standard evaluation data set;

[0022] Determine the weights of each evaluation indicator in the evaluation indicator system of various effective facilities by the entropy weight method;

[0023] According to the weights and the standard evaluation data set, calculate the comprehensive evaluation values of various effective facilities by the linear weighted method to obtain a comprehensive evaluation data set of effective facilities; wherein, the comprehensive evaluation data set at least includes: the longitude and latitude coordinates, the comprehensive evaluation value and the action category of the effective facilities.

[0024] In a feasible implementation manner, obtain the boundary longitude and latitude data set of the area to be selected, create a vector map of the area to be selected, and perform grid processing on the vector map of the area to be selected to obtain a grid distribution map, specifically including:

[0025] Obtain the boundary longitude and latitude data of all administrative regions involved in the area to be selected to obtain the boundary longitude and latitude data set of the area to be selected;

[0026] If there are incomplete administrative regions involved in the area to be selected, delimit a polygonal broken line area including the area to be selected in the incomplete administrative region, and determine the boundary longitude and latitude data set of the area to be selected through the longitude and latitude coordinates of the polygonal broken line area;

[0027] Based on the boundary longitude and latitude data set, create a vector map of the area to be selected;

[0028] According to the map information in the vector map of the area to be selected, the coverage radius of the facility to be selected, and the sampling address point sampling density requirements, set the type, creation range, horizontal interval and vertical interval of the grid area, and select the same reference coordinate system as the vector map of the area to be selected to create a grid for the vector map of the area to be selected to obtain the grid distribution map; wherein, the map information in the vector map of the area to be selected includes: the range and area of the area to be selected, and the association rule between longitude and latitude and distance in the area to be selected.

[0029] In a feasible implementation manner, obtain a sampling address data set in the grid distribution map through an address uniform sampling mechanism, specifically including:

[0030] According to the address uniform sampling mechanism, determine the basic sampling lattice points corresponding to each grid unit in the grid distribution map; the basic sampling lattice points include 1 central sampling address point and 4 corresponding associated sampling address points, and the digital marks corresponding to each sampling address point;

[0031] Calculate the longitude and latitude coordinates of the centroid of each grid cell based on the longitude and latitude data of the grid boundary lines, and obtain the longitude and latitude coordinates of the central sampling address point;

[0032] Calculate the longitude and latitude coordinates of the associated sampling address points according to the association rule between the central sampling point and the associated sampling address points; and determine the digital marks of each sampling address point according to the digital marking generation rule, so as to obtain the sampling address data set; the sampling address data set includes: the longitude and latitude coordinates and digital marks of each sampling address point.

[0033] In a feasible implementation manner, based on the comprehensive evaluation data set and the sampling address data set, determine the sampling address annotation data set, specifically including:

[0034] Based on the comprehensive evaluation data set and the sampling address data set, use the discount quantization function of the comprehensive evaluation value of the effective facilities: Calculate the discount coefficient φ(p, q) corresponding to the comprehensive evaluation value of the effective facilities for each sampling address point; where r is the coverage radius of the facility to be located; λ ∈ (0, 1) is the proportion of the group served within half of the coverage radius to the total group served; d(p, q) is the distance between the sampling address point p and the associated effective facility q;

[0035] According to Calculate the positive decision association data of the sampling address point where p is the sampling address point, q + is the positive effective facility, φ(p, q + ) is the discount coefficient corresponding to the comprehensive evaluation value of the effective facility q + for the sampling address point p, is the comprehensive evaluation value of the effective facility q + ;

[0036] According to Calculate the negative decision association data of the sampling address point where p is the sampling address point, q - is the negative effective facility, φ(p, q - ) is the discount coefficient corresponding to the comprehensive evaluation value of the effective facility q - for the sampling address point p, is the comprehensive evaluation value of the effective facility q - ;

[0037] Synchronize the obtained positive decision - associated data and negative decision - associated data to the sampling address data set to obtain a sampling address decision - associated data set; at least included in the sampling address decision - associated data set are: the longitude and latitude coordinates of the sampling address point, digital markers, positive decision - associated data, and negative decision - associated data;

[0038] Determine the label annotation method for the sampling address point, and use the annotation method to perform label annotation on the sampling address decision - associated data set to obtain the sampling address annotation data set; wherein, at least included in the sampling address annotation data set are: the longitude and latitude coordinates of the sampling address point, marked numbers, and point location labels.

[0039] In a feasible implementation manner, determining the label annotation method for the sampling address point and using the annotation method to perform label annotation on the sampling address decision - associated data set to obtain the sampling address annotation data set specifically includes:

[0040] According to the site - selection decision feedback form, determine the label form of the sampling address point; wherein, the site - selection decision feedback form includes a scoring feedback form and a rating feedback form; the label form includes a scoring label form and a rating label form;

[0041] If the determined label form is a scoring label form, then based on a scoring function, calculate the scoring data of each sampling address point in the sampling address decision - associated data set, and label the scoring data as the point location label of the sampling address point;

[0042] If the determined label form is a rating label form, then based on a weighting function, calculate the weighted decision - associated data of each sampling address point in the sampling address decision - associated data set; based on a rating function, map the weighted decision - associated data to the corresponding rating data, and label the rating data as the point location label of the sampling address point.

[0043] In a feasible implementation manner, according to the distribution law of the point location labels of the sampling address points in the sampling address annotation data set, select the corresponding auxiliary site - selection model, specifically including:

[0044] Analyze the distribution law of the point location labels of each sampling address point in the sampling address annotation data set, determine that the distribution law of the point location labels is a non - linear distribution, and preliminarily select an auxiliary site - selection model of the non - linear model type;

[0045] If the sampling address point label is in the scoring label form, the selected auxiliary site - selection model is a non - linear regression model; if the sampling address point label is in the rating label form, the selected auxiliary site - selection model is a non - linear classification model.

[0046] In a feasible implementation, the auxiliary site selection model is trained and tuned by annotating the dataset with the sampling addresses to obtain the final auxiliary site selection model, which specifically includes:

[0047] According to the digital tags of each sampling address point in the sampling address-annotated dataset, determine the splitting scheme of the sampling address-annotated dataset, and split the sampling address-annotated dataset into a training dataset, a validation dataset, and a test dataset;

[0048] Through the training dataset and the validation dataset, perform hyperparameter tuning on the auxiliary site selection model to determine the hyperparameter combination with the optimal performance;

[0049] Merge the training dataset and the validation dataset into a new training dataset;

[0050] Based on the new training dataset and the hyperparameter combination, perform model training on the auxiliary site selection model to obtain the final auxiliary site selection model;

[0051] Based on the test dataset, calculate the evaluation metrics of the final auxiliary site selection model to test and evaluate the performance of the final auxiliary site selection model.

[0052] On the other hand, an embodiment of the present invention also provides an auxiliary site selection device based on machine learning. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, so that the at least one processor can execute the auxiliary site selection method based on machine learning.

[0053] Compared with the prior art, an auxiliary site selection method and device based on machine learning provided by the present invention have the following beneficial effects:

[0054] After determining the annotated address dataset based on the data involved in the site selection decision, the present invention uses a machine learning algorithm to learn the rules and patterns from the annotated address dataset, obtains an auxiliary site selection model with high generalization performance, and uses the model to give real-time auxiliary site selection decision results according to the spatial location characteristics of the addresses to be selected, providing reasonable decision guidance and effective technical support for the site selection planning. By optimizing the auxiliary site selection decision process and mining the patterns and relationships between the data involved in the auxiliary site selection, the present invention can solve the problem of the long time-consuming real-time decision-making process of the auxiliary site selection caused by the reasons such as the variety, large scale, and many processing processes of the data involved in the auxiliary site selection. Description of the Drawings

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0056] Figure 1 Flowchart of an auxiliary site selection method based on machine learning provided by an embodiment of the present invention;

[0057] Figure 2 Flowchart of uniform sampling of sampling address points provided by an embodiment of the present invention;

[0058] Figure 3 Flowchart of label marking of sampling address points provided by an embodiment of the present invention;

[0059] Figure 4 Flowchart of training, optimization and evaluation of an auxiliary site selection model provided by an embodiment of the present invention;

[0060] Figure 5 Schematic structural diagram of an auxiliary site selection device based on machine learning provided by an embodiment of the present invention. Detailed implementation manners

[0061] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0062] An embodiment of the present invention provides an auxiliary site selection method based on machine learning, as Figure 1 shown. The auxiliary site selection method based on machine learning specifically includes steps S101 - S107:

[0063] S101. Based on the site selection requirements, determine various types of effective facilities within the area to be selected, and construct an evaluation index system for the effective facilities.

[0064] Specifically, first, based on the site selection requirements and the usage nature of the facilities to be selected, determine several types of effective facilities that affect the site selection decision; among them, the several types of effective facilities at least include: residential providing facilities, medical service facilities, educational service facilities, shopping service facilities, and transportation service facilities.

[0065] Further, according to the relationship between the effective facilities and the facilities to be located, determine the type of influence of the effective facilities on the location decision-making; wherein, the relationship between the effective facilities and the facilities to be located includes a dependency relationship and a competition relationship; the type of influence includes a positive influence and a negative influence.

[0066] Further, according to the type of influence, determine the role category of the effective facilities; wherein, the role category of the effective facilities corresponding to the positive influence is the positive effective facilities, and the role category of the effective facilities corresponding to the negative influence is the negative effective facilities.

[0067] Further, based on the facility attribute characteristics of the effective facilities of each use type, construct an evaluation index system corresponding to the effective facilities of each use type; wherein, the evaluation indexes in the evaluation index system include: relevant indexes of the population living inside the facilities, relevant indexes of the population receiving services inside the facilities, relevant indexes of the service content of the facilities, and relevant indexes of the internal attributes of the facilities.

[0068] As a feasible implementation method, after obtaining the location selection requirements, first determine the effective facilities of various use types that affect the location decision-making and divide the corresponding role categories, and construct an evaluation index system corresponding to the effective facilities. The specific implementation method is as follows:

[0069] (1) During the location decision-making process, it is necessary to combine the actual requirements of the location decision-making and the use nature of the facilities to be located to determine the effective facilities of different use types in the area to be located that can affect the location decision-making. The types of effective facilities specifically include: residential providing facilities (residential buildings, communities or neighborhoods), medical service facilities (hospitals and medical service institutions of a certain scale), education service facilities (kindergartens, primary schools, middle schools, universities), shopping service facilities (shopping malls, supermarkets, markets of a certain scale), transportation service facilities (subway stations, bus stops), etc.

[0070] After determining the effective facilities of various uses considered in the location decision-making, it is necessary to divide the effective facilities into effective facilities with different influencing effects based on the relationship between the effective facilities of each use and the location facilities. The specific division method is as follows: If the relationship between the effective facilities of a certain use type and the location facilities is a dependency relationship, that is, the effective facilities of this type or the group among them need the resources and services provided by the location facilities, then the effective facilities of this type have a positive influence on the location decision-making of the location facilities, and the effective facilities of this type can be divided into positive effective facilities; if the relationship between the effective facilities of a certain use type and the location facilities is a competition relationship, that is, the resources and services that the effective facilities of this type can provide are the same or similar to those that the location facilities can provide, then the effective facilities of this type have a negative influence on the location decision-making of the location facilities, and the effective facilities of this type can be divided into negative effective facilities.

[0071] (2) In order to comprehensively and quantitatively evaluate the demand scale or service scale of effective facility entities, it is necessary to combine the location selection influencing factors and the facility attribute characteristics to determine the evaluation index system corresponding to the effective facilities of each use type. The evaluation indicators in the evaluation index system specifically include: various indicators related to the population living inside the facility of this type, various indicators related to the population receiving services inside the facility of this type, various indicators related to the service content of the facility of this type, and various indicators related to the internal attributes of the facility of this type.

[0072] S102. Normalize the evaluation data of the same type of effective facilities to obtain the comprehensive evaluation data set corresponding to each type of effective facilities.

[0073] Specifically, obtain the spatial location data and evaluation data of various effective facilities in the area to be selected; among them, the spatial location data includes at least the longitude and latitude of the effective facilities; the evaluation data includes at least the evaluation index data of the effective facilities in their evaluation index system.

[0074] Furthermore, divide the evaluation indicators in the evaluation index system into positive evaluation indicators and negative evaluation indicators. Then, based on the linear transformation algorithm, perform linear transformation on the evaluation data corresponding to the positive evaluation indicators and the evaluation data corresponding to the negative evaluation indicators respectively, and convert the evaluation data into standard evaluation data within the range of [0, 1] with the same action direction to obtain the standard evaluation data set.

[0075] Furthermore, determine the weights of the evaluation indicators in the evaluation index system of various effective facilities through the entropy weight method. According to the weights and the standard evaluation data set, calculate the comprehensive evaluation values of various effective facilities through the linear weighted method to obtain the comprehensive evaluation data set of the effective facilities. Among them, the comprehensive evaluation data set includes at least: the longitude and latitude coordinates of the effective facilities, the comprehensive evaluation value, and the action category.

[0076] As a feasible implementation method, after constructing the evaluation index system corresponding to the effective facilities of each use type, it is necessary to obtain the spatial location data and various evaluation data of each type of effective facility entity in the area to be selected. The spatial location data mainly includes the longitude and latitude of the facility entity, which is used for the positioning of the facility entity and the calculation of the distance between the facility entity and the address sampling point; the evaluation data mainly includes the evaluation index data of the facility entity in its corresponding evaluation index system, which is used for the description of the specific situation of the facility entity and the comprehensive quantitative evaluation of the demand scale or service scale of the facility entity.

[0077] Since there may be certain differences in the properties, dimensions, orders of magnitude, etc. of the evaluation indicators in the evaluation index system corresponding to the effective facilities of each use type, in order to eliminate the influence of dimensions, unify the comparison criteria, ensure the consistency of data properties and the reliability of evaluation results, it is necessary to normalize the evaluation data of each effective facility entity of the same use type to obtain the standard evaluation data set corresponding to the effective facilities of each use type. The specific processing process includes: according to the different effects of the evaluation indicators on the evaluation results, the indicators in the evaluation index system with larger values making the evaluation results better are marked as positive evaluation indicators, and the indicators with smaller values making the evaluation results better are marked as negative evaluation indicators. Then, based on the linear proportion transformation method or the range transformation method, corresponding linear transformations are performed on the positive and negative evaluation indicators to convert the original evaluation data into standard evaluation data between 0 and 1, with the same direction of action and comparability, so as to obtain the standard evaluation data set corresponding to the effective facilities of each use type.

[0078] Then, based on the standard evaluation data set of the effective facilities, the entropy weight method is used to determine the weights of the evaluation indicators in the evaluation index system corresponding to the effective facilities of each use type, and the linear weighted method is used to calculate the comprehensive evaluation value of the effective facility entity. The specific implementation method is as follows:

[0079] For the evaluation index system corresponding to the effective facilities of each use type, the weights of the evaluation indicators can be determined based on the standard evaluation data corresponding to the effective facility entities of the same use type using the entropy weight method. The entropy weight method is an objective weighting method, which mainly determines the weights of the indicators according to the dispersion degree of the actual evaluation data of the evaluation indicators. Specifically, for the evaluation indicators in the evaluation index system, calculate the proportion value of the evaluation data of the effective facility entity corresponding to this indicator in the total sum of all evaluation data of this indicator, and calculate the information entropy of this indicator based on this proportion value. Finally, calculate the information entropy redundancy of the evaluation indicator to determine the weight of the evaluation indicator.

[0080] After determining the weights of the evaluation indicators in the evaluation index system corresponding to the effective facilities of each use type, the linear weighted method can be used in combination with the standard evaluation data set corresponding to the effective facilities of each use type to calculate the comprehensive evaluation value of the effective facility entity. The specific process: for the effective facility entity of each use type, perform weighted averaging on its corresponding multiple standard evaluation data to obtain a comprehensive evaluation value that can evaluate the quality of the facility entity, and then an effective facility comprehensive evaluation data set can be obtained. This data set mainly includes the longitude and latitude coordinates, comprehensive evaluation value, and action category (including positive effective facilities and negative effective facilities) of the facility entity.

[0081] S103. Obtain the boundary longitude and latitude data set of the area to be selected for site, create a vector map of the area to be selected for site, and perform grid processing on the vector map of the area to be selected for site to obtain a grid distribution map.

[0082] Specifically, obtain the boundary longitude and latitude data of all administrative regions involved in the area to be selected, and obtain the boundary longitude and latitude data set of the area to be selected. If there are incomplete administrative regions involved in the area to be selected, a polygonal broken line area containing the area to be selected is delimited in the incomplete administrative region, and the boundary longitude and latitude data set of the area to be selected is determined through the longitude and latitude coordinates of the polygonal broken line area.

[0083] Further, based on the boundary longitude and latitude data set, create a vector map of the area to be selected. Then, according to the map information in the vector map of the area to be selected, the coverage radius of the facility to be selected, and the sampling density requirements of the sampling address points, set the type, creation range, horizontal interval, and vertical interval of the grid area, and select the same reference coordinate system as the vector map of the area to be selected to create a grid for the vector map of the area to be selected, and obtain a grid distribution map. Among them, the map information in the vector map of the area to be selected includes: the range and area of the area to be selected, and the association rule between longitude and latitude and distance within the area to be selected.

[0084] S104. Obtain a sampling address data set in the grid distribution map through an address uniform sampling mechanism.

[0085] Specifically, according to the address uniform sampling mechanism, determine the basic sampling dot matrix corresponding to each grid cell in the grid distribution map; the basic sampling dot matrix includes 1 central sampling address point and corresponding 4 associated sampling address points, and a digital label corresponding to each sampling address point.

[0086] Further, according to the longitude and latitude data of the grid side line, calculate the longitude and latitude coordinates of the centroid of each grid cell to obtain the longitude and latitude coordinates of the central sampling address point. According to the association rule between the longitude and latitude coordinates of the central sampling point and the associated sampling address points, calculate the longitude and latitude coordinates of the associated sampling address points; and according to the digital label generation rule, determine the digital labels of each sampling address point to obtain a sampling address data set; the sampling address data set includes: the longitude and latitude coordinates and digital labels of each sampling address point.

[0087] As a feasible implementation manner, Figure 2 For the uniform sampling flow chart of sampling address points provided by the embodiments of the present invention, as Figure 2 shown, the specific implementation manner of obtaining the sampling address data set is as follows:

[0088] (1) First, by analyzing the specific composition of the area to be selected, obtain the longitude and latitude data set of the boundary of the area to be selected, create a vector map of the area to be selected, and grid the vector map to obtain a grid distribution map. The specific process is as follows: if the area to be selected is composed of a single or multiple administrative divisions, it is necessary to obtain the longitude and latitude data sets of the boundaries of all administrative divisions involved, and obtain the longitude and latitude data set of the boundary of the area to be selected by combining these data sets; if the area to be selected involves incomplete administrative divisions, it is necessary to delineate the polygonal polyline area containing the area to be selected, and obtain the longitude and latitude data set of the boundary of the area to be selected by combining the longitude and latitude coordinates of these polylines. The acquired longitude and latitude dataset of the boundary of the area to be selected is converted into a file in vector graphics format. Based on the geometric position and related attribute information stored in the file, the geographic information system software is used to realize the visualization of the longitude and latitude dataset of the boundary of the area to be selected, and a vector map of the area to be selected is created; the type and creation range of the grid area are determined according to the overall scope of the area to be selected and the coverage radius of the site selection facilities; the horizontal and vertical intervals of the grid area are determined according to the area of ​​the area to be selected and the requirements of the sampling density, combined with the correlation between longitude and latitude and distance within the scope of the area to be selected; according to the basic attributes of the vector map of the area to be selected, the reference coordinate system of the grid is selected in accordance with the vector map. Figure 1 After completing the basic parameter setting of the grid, use the geographic information system software to create a grid for the vector map of the site selection area to obtain the initial grid distribution map, and then use an appropriate method to remove the grids that are not within the site selection area according to the site selection range to obtain the grid distribution map.

[0089] In one embodiment, the longitude and latitude data set of the designated administrative division boundary, or the demarcated polygonal polyline area, can be obtained through map platforms such as DataV and Tiandi Map, and exported as a longitude and latitude data set of the boundary of the area to be selected in geojson format. The longitude and latitude data set of the boundary of the area to be selected is converted into a vector data file of the boundary of the area to be selected in shp format using a geographic data conversion algorithm or a converter. The vector data file of the boundary of the area to be selected is imported using the geographic information system software QGIS to create a vector layer, and a vector map of the area to be selected is obtained. The layer where the vector map of the area to be selected is located is used as the bottom layer, and a vector creation tool is used to create a grid layer in a suitable reference coordinate system according to the grid parameters such as the shape, creation range, horizontal interval, and vertical interval of the set grid to obtain an initial grid distribution map. The layer where the initial grid distribution map is located is used as the input layer, and the layer where the vector map of the area to be selected is used as the overlay layer. The vector geoprocessing tool is used to obtain a grid layer superimposed with the vector map of the area to be selected by clipping, and a grid distribution map of the area to be selected is obtained.

[0090] (2) After obtaining the grid distribution map, design an address uniform sampling mechanism according to the grid distribution law to determine the longitude and latitude coordinate data of the sampling points and the usage types of the sampling points. The specific content of the address uniform sampling mechanism is as follows: Based on the grid distribution map, calculate the longitude and latitude coordinates of the centroid of each grid cell (C lon ,C lat ) according to the longitude and latitude data of the grid edges, and determine the grid centroid as the central sampling address point. For each central sampling address point, determine 4 different associated sampling address points corresponding to it, namely sampling address point A: (C lon +α,C lat +β), sampling address point B: (C lon -α,C lat +β), sampling address point C: (C lon -α,C lat -β), sampling address point D: (C lon +α,C lat -β); where α and β are both fixed parameters, which can be determined by the grid attribute parameters and the correlation law between longitude, latitude and distance: α makes and be separated by a quarter of the horizontal grid interval length; β makes and be separated by a quarter of the vertical grid interval length; and are the minimum longitude, maximum longitude, minimum latitude and maximum latitude of all central sampling address points respectively. For a single grid cell, the type of the central sampling address point included in this grid cell is marked as the number 5, and the 4 associated sampling points corresponding to the central sampling point need to be randomly marked with different integers from 1 to 4. The 5 sampling address points and the corresponding marking information included in a single grid cell are used as a basic sampling dot matrix. It should be noted that the sampling address points generated based on the address uniform sampling mechanism all maintain a reasonable distance interval, ensuring the uniformity of sampling.

[0091] According to the above address uniform sampling mechanism, use geographic information system software to determine the longitude and latitude coordinates of the central sampling address points included in each basic sampling dot matrix in the grid distribution map, and obtain the longitude and latitude coordinate set of the central sampling address points. Then, based on the longitude and latitude coordinate set of the central sampling address points, use a data processing tool to calculate the longitude and latitude coordinates of the associated sampling address points included in each basic sampling dot matrix, and obtain the longitude and latitude coordinate set of the associated sampling address points. Based on the longitude and latitude coordinate set of the central sampling address points and the longitude and latitude coordinate set of the associated sampling address points, mark digital marks on the sampling address points in the basic sampling dot matrix to obtain a sampling address data set, which mainly includes the longitude, latitude and digital marks of the sampling address points.

[0092] In one embodiment, based on the grid distribution map, the longitude and latitude coordinates of the centroid of each grid cell are determined using the QGIS field calculator tool of geographic information system software based on the longitude and latitude data of the grid edges, and the obtained dataset is exported as the longitude and latitude dataset of the central sampling address points; the data processing tool Python is used to read the longitude and latitude dataset of the central sampling address points, calculate the longitude and latitude coordinates of the associated sampling address points according to the association relationship between the associated sampling address points and the central sampling address points in the address uniform sampling mechanism, and randomly mark the sampling points with the set numbers.

[0093] S105. Determine the sampling address annotation dataset based on the comprehensive evaluation dataset and the sampling address dataset.

[0094] Specifically, based on the comprehensive evaluation dataset and the sampling address dataset, use the discount quantization function of the comprehensive evaluation value of the effective facilities: Calculate the discount coefficient φ(p, q) of the comprehensive evaluation value of the effective facilities corresponding to each sampling address point; where r is the coverage radius of the facility to be located; λ ∈ (0, 1) is the proportion of the population served within half of the coverage radius to the total population served; d(p, q) is the distance between the sampling address point p and the associated effective facility q.

[0095] Further, according to Calculate the positive decision association data of the sampling address point where p is the sampling address point, q + is the positive effective facility, φ(p, q + ) is the discount coefficient of the comprehensive evaluation value of the effective facility q + corresponding to the sampling address point p, and v q+ is the comprehensive evaluation value of the effective facility q + .

[0096] Further, according to Calculate the negative decision association data of the sampling address point where p is the sampling address point, q - is the negative effective facility, φ(p, q - ) is the discount coefficient of the comprehensive evaluation value of the effective facility q - corresponding to the sampling address point p, and v q- is the comprehensive evaluation value of the effective facility q - .

[0097] Then, synchronize the obtained positive decision - associated data and negative decision - associated data to the sampling address data set to obtain a sampling address decision - associated data set. The sampling address decision - associated data set at least includes: the longitude and latitude coordinates of the sampling address point, digital markers, positive decision - associated data, and negative decision - associated data.

[0098] Furthermore, determine the label annotation method for the sampling address points, and use the annotation method to perform label annotation on the sampling address decision - associated data set to obtain a sampling address annotation data set. Among them, the sampling address annotation data set at least includes: the longitude and latitude coordinates of the sampling address point, marker numbers, and point location labels.

[0099] Then, according to the site - selection decision feedback form, determine the label form of the sampling address points; among them, the site - selection decision feedback form includes a scoring feedback form and a rating feedback form; the label form includes a scoring label form and a rating label form.

[0100] If the determined label form is the scoring label form, then based on the scoring function, calculate the scoring data of each sampling address point in the sampling address decision - associated data set, and label the data as the point location label of the sampling address point; if the determined label form is the rating label form, then based on the weighting function, calculate the weighted decision - associated data of each sampling address point in the sampling address decision - associated data set; based on the rating function, map the weighted decision - associated data to the corresponding rating data, and label the rating data as the point location label of the sampling address point.

[0101] As a feasible implementation method, Figure 3 For the sampling address point label marking flow chart provided by the embodiment of the present invention, as Figure 3 shown, the specific steps to determine the sampling address annotation data set are as follows:

[0102] (1) First, construct a discount quantization function for the comprehensive evaluation value of effective facilities.

[0103] During the site - selection decision - making process, the distance between the site - selection point and the effective facility entity will affect the site - selection decision result. For positive effective facilities, the closer it is to the site - selection point, the greater its positive impact on the site - selection decision result; for negative effective facilities, the closer it is to the site - selection point, the greater its negative impact on the site - selection decision result. Considering the dynamic influence of the distance factor on the site - selection decision, in order to describe the influence law of the distance between the site - selection point and the effective facility on the decision, it is necessary to construct a discount quantization function for the comprehensive evaluation value of effective facilities to quantify the influence degree of the distance between the site - selection point and the effective facility on the decision. The discount quantization function of the comprehensive evaluation value of effective facilities can calculate the discount coefficient corresponding to the address point of the comprehensive evaluation value of the facility entity based on the spatial position data of the effective facility and the site - selection point. The specific form of the discount quantization function of the comprehensive evaluation value of effective facilities is:

[0104] Where r is the coverage radius of the facility to be located, which can be determined by considering factors such as the resource and service provision capabilities of the facility to be located, the ways in which effective facilities or the groups within them obtain resources and services, etc.; λ ∈ (0, 1) is the proportion of the group receiving services within half of the effective coverage radius of the facility to be located to the total group receiving services; d(p, q) is the distance between the location point p and the associated effective facility q, which can be calculated using the haversine formula based on the longitudes and latitudes of the location point p and the associated effective facility q.

[0105] (2) Combine the comprehensive evaluation dataset of effective facilities and the discount quantization function of the comprehensive evaluation value to obtain the decision - associated dataset for sampling addresses.

[0106] Based on the sampling address dataset and the comprehensive evaluation dataset of effective facilities, for a given sampling address point, use the discount quantization function of the comprehensive evaluation value of effective facilities to calculate the discount coefficient corresponding to the comprehensive evaluation value of the effective facility entity for this sampling address point, and further calculate the decision - associated data for the sampling address point according to the function category of the effective facility, mainly including positive decision - associated data and negative decision - associated data.

[0107] The positive decision - associated data of the sampling address point is the sum of the discount values corresponding to the comprehensive evaluation values of all positive facility entities in the comprehensive evaluation dataset of effective facilities for this sampling address point. The specific calculation method is: Where p is the sampling address point, q + is a positive facility entity, and v q+ is the comprehensive evaluation value of the facility entity q + ;

[0108] The negative decision - associated data of the sampling address point is the sum of the discount values corresponding to the comprehensive evaluation values of all negative facility entities in the comprehensive evaluation dataset of effective facilities for this sampling address point. The specific calculation method is: Where p is the sampling address point, q - is a negative facility entity, and v q- is the comprehensive evaluation value of the facility entity q - ;

[0109] Finally, synchronize the obtained positive decision - associated data and negative decision - associated data to the sampling address dataset to obtain the decision - associated dataset for sampling addresses. This dataset mainly includes the longitude, latitude, digital label, positive decision - associated data, and negative decision - associated data of the sampling address point.

[0110] (3) Determine the form of the label according to the form of the location decision feedback, determine the label marking method for the sampling address point, and obtain the labeled dataset for sampling addresses based on the decision - associated dataset for sampling addresses.

[0111] The feedback form of the site selection decision mainly considers the form of scoring or rating, and the specific feedback form selected can be determined according to actual needs. The feedback form of the site selection decision will affect the form of the labels of the sampled address points. If the scoring feedback form is selected, the labels of the sampled address points are continuous score values; if the rating feedback form is selected, the labels of the sampled address points are discrete rating categories.

[0112] For the scoring label form, based on the decision - associated data set of the sampled address, the scoring data of the sampled address points are calculated by constructing a scoring function according to the positive decision - associated data and negative decision - associated data corresponding to the sampled address points, and the scoring data are labeled as the labels of the sampled address points to obtain the sampled address annotation data set. The scoring function of the sampled address points is specifically:

[0113]

[0114] where and represent the positive decision - associated data and negative decision - associated data respectively, and represent the maximum values of the positive decision - associated data and negative decision - associated data respectively, and w + and w - are the weights of the positive decision - associated data and negative decision - associated data respectively.

[0115] For the rating label form, based on the decision - associated data set of the sampled address, the weighted decision - associated data of the sampled address points are calculated by constructing a weighting function according to the positive decision - associated data and negative decision - associated data corresponding to the sampled address points. The rating parameters are determined according to the overall distribution of the weighted decision - associated data, and then a rating function is constructed based on the rating parameters to map the weighted decision - associated data of the sampled address points to the corresponding ratings.

[0116] The weighted function of the decision - associated data of the sampled address points is specifically: where p is the sampled address point, and represent the positive decision - associated data and negative decision - associated data respectively, represents the maximum value of the negative decision - associated data, and w + and w - are the weights of the positive decision - associated data and negative decision - associated data respectively.

[0117] The rating function of the sampled address points is specifically:

[0118]

[0119] where p is the sampled address point, m 1 、m 2 、m 3For rating parameters.

[0120] The method for determining the rating parameter is as follows: Consider all sampling address points, and calculate the median of the weighted decision - associated data as m. 3 ; Consider all sampling address points where the weighted decision - associated data is not less than m. 3 Calculate the median of the weighted decision - associated data as m. 2 ; Consider all sampling address points where the weighted decision - associated data is not less than m. 2 Calculate the median of the weighted decision - associated data as m. 1 .

[0121] Determine the form of the sampling address point label according to the form of the siting decision feedback, and use the generation method with the corresponding label form based on the sampling address decision - associated data set to determine the label of the sampling address point, obtaining the labeled address data set, which mainly includes the longitude, latitude, marked number, and point label of the sampling address point.

[0122] S106. According to the distribution law of the point labels of the sampling address points in the sampling address labeled data set, select the corresponding auxiliary siting model.

[0123] Specifically, analyze the distribution law of the point labels of each sampling address point in the sampling address labeled data set, determine that the distribution law of the point labels is a non - linear distribution, and initially select an auxiliary siting model of the non - linear model type. If the sampling address point label is in the form of a scoring label, the selected auxiliary siting model is a non - linear regression model; if the sampling address point label is in the form of a rating label, the selected auxiliary siting model is a non - linear classification model.

[0124] S107. Through the sampling address labeled data set, train the auxiliary siting model and optimize it to obtain the final auxiliary siting model for making auxiliary siting decisions through the final auxiliary siting model.

[0125] Specifically, according to the digital markings of each sampling address point in the sampling address labeled data set, determine the splitting scheme of the sampling address labeled data set, and split the sampling address labeled data set into a training data set, a validation data set, and a test data set.

[0126] Furthermore, through the training data set and the validation data set, perform hyperparameter tuning on the auxiliary siting model to determine the hyperparameter combination with the optimal performance. Merge the training data set and the validation data set into a new training data set. Based on the new training data set and the hyperparameter combination, perform model training on the auxiliary siting model to obtain the final auxiliary siting model.

[0127] Further, based on the test data set, evaluate the evaluation metrics of the final auxiliary site selection model to test and evaluate the performance of the final auxiliary site selection model. After passing the evaluation, the final auxiliary site selection model can be put into use to generate auxiliary site selection decisions.

[0128] As a feasible implementation manner, Figure 4 The flowchart of training, optimizing, and evaluating an auxiliary site selection model provided by an embodiment of the present invention is as Figure 4 shown. The steps for obtaining the auxiliary site selection model are as follows:

[0129] (1) After determining the labeled address data set, the longitude and latitude of the sampled address points in this data set are used as the feature data for the auxiliary site selection model to learn, the marked numbers are used as the filtering data for splitting the data set, and the point labels are used as the label data for the auxiliary site selection model to learn. Based on this, determine the splitting scheme of the sampled address labeled data set according to the digital marks of the sampled address points, and split this data set into a training data set, a validation data set, and a test data set. The splitting scheme of the sampled address labeled data set is as follows: all the feature data and label data corresponding to the sampled address points with digital marks less than 4 are screened out as the training data set; all the feature data and label data corresponding to the sampled address points with digital marks of 5 are screened out as the validation data set; all the feature data and label data corresponding to the sampled points with digital marks of 4 are screened out as the test data set. Due to the uniformity of the digital marks, the above-mentioned splitting scheme of the labeled address data set can ensure the rationality of splitting the labeled address data set and maintain the data ratio of the training data set, the validation data set, and the test data set as 3:1:1.

[0130] Then, according to the form and distribution characteristics of the address labels, determine the type and training algorithm of the auxiliary site selection model. By analyzing the distribution characteristics of the labels of the sampled address points, it can be found that the auxiliary site selection problem has obvious non-linear characteristics. Therefore, the type of the auxiliary site selection model should be a non-linear model. If the label of the sampled address point is in the form of a score, the type of the auxiliary site selection model should be a non-linear regression model, including models such as polynomial regression, support vector regression, and neural network; if the label of the sampled address point is in the form of a rating, the type of the auxiliary site selection model should be a non-linear classification model, including models such as support vector machine, random forest, and neural network.

[0131] (2) Based on the data set after dividing the labeled address data set, determine the hyperparameter combination that makes the performance of the auxiliary site selection model optimal through hyperparameter tuning, and use the training algorithm corresponding to the auxiliary site selection model to determine the final auxiliary site selection model.

[0132] Specifically, based on the type of the auxiliary site selection model and the corresponding training algorithm, the types and value ranges of hyperparameters are set to determine the hyperparameter combination search space. Hyperparameter tuning is a process of finding the hyperparameter combination that optimizes the overall performance of the auxiliary site selection model based on the hyperparameter combination search space. The process of hyperparameter tuning includes: after selecting a set of hyperparameter combinations from the hyperparameter combination search space, based on the training data set, using the training algorithm corresponding to the auxiliary site selection model to perform model training to obtain an initial auxiliary site selection model; based on the validation data set, calculating and recording the evaluation metric values that can evaluate the overall performance of the initial auxiliary site selection model. Among them, if the initial auxiliary site selection model belongs to a non-linear regression model, the corresponding evaluation metrics include absolute error, mean square error, goodness of fit, etc.; if the initial auxiliary site selection model belongs to a non-linear classification model, the evaluation metrics include accuracy, precision, recall, etc. By comparing the recorded evaluation metric values of the initial auxiliary site selection model, the hyperparameter combination corresponding to the initial auxiliary site selection model with the optimal overall performance is determined.

[0133] After determining the optimal hyperparameter combination through the hyperparameter tuning process, the training data set and the validation data set are merged into a new training data set; based on the merged training data set, using the training algorithm corresponding to the auxiliary site selection model to perform model training to obtain a final auxiliary site selection model. Finally, based on the test data set, the evaluation metrics of the final auxiliary site selection model are calculated to test and evaluate the performance of the final auxiliary site selection model.

[0134] In one embodiment, the Scikit-learn machine learning library can be used to complete the training, optimization, and evaluation of the auxiliary site selection model. First, load the training data set, validation data set, and test data set, and determine a suitable model training algorithm in Scikit-learn according to the type of the auxiliary site selection model; then, use the grid search method in Scikit-learn to complete hyperparameter tuning according to the requirements of the hyperparameter tuning process to determine the hyperparameter combination that optimizes the overall performance of the auxiliary site selection model in the hyperparameter combination search space; finally, based on the merged training data set, use the model training algorithm to perform model training to obtain a final auxiliary site selection model, and use the test data set to test and evaluate this model.

[0135] In addition, an embodiment of the present invention also provides an auxiliary site selection device based on machine learning, as Figure 5 shown, the auxiliary site selection device based on machine learning specifically includes:

[0136] At least one processor; and, a memory communicatively connected to the at least one processor; wherein,

[0137] The memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute:

[0138] Based on the site selection requirements, various effective facilities within the area to be selected are determined, and an evaluation index system for effective facilities is constructed;

[0139] Normalize the evaluation data of the same type of effective facilities to obtain a comprehensive evaluation data set corresponding to each type of effective facility;

[0140] Obtain the boundary longitude and latitude data set of the area to be selected, create a vector map of the area to be selected, and perform grid processing on the vector map of the area to be selected to obtain a grid distribution map;

[0141] Through an address uniform sampling mechanism, obtain a sampling address data set in the grid distribution map;

[0142] Based on the comprehensive evaluation data set and the sampling address data set, determine a sampling address annotation data set;

[0143] According to the distribution law of the point position labels of the sampling address points in the sampling address annotation data set, select a corresponding auxiliary site selection model;

[0144] Through the sampling address annotation data set, train and optimize the auxiliary site selection model to obtain a final auxiliary site selection model for making auxiliary site selection decisions through the final auxiliary site selection model.

[0145] Machine learning is a technology that, based on a large amount of effective data, uses specific algorithms to enable a computer to automatically learn, recognize, and reason, and continuously optimize and improve its processing ability. During the machine learning process, the computer can automatically discover the laws and patterns in the data through the analysis and learning of the data, and make predictions and decisions based on these laws and patterns. Based on the characteristics of the complexity of site selection decisions, the present invention optimizes the site selection decision-making process in combination with the characteristics of machine learning technology, simplifies the calculation method of decision-making results, and has high practical value.

[0146] Each embodiment in the present invention is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0147] The above describes specific embodiments of the present invention. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0148] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the present invention.

Claims

1. A machine learning-based auxiliary site selection method, characterized in that, the method includes: Based on the site selection requirements, determine various types of effective facilities within the area to be selected, and construct an evaluation index system for effective facilities; Normalize the evaluation data of the same type of effective facilities to obtain a comprehensive evaluation data set corresponding to each type of effective facility; Obtain the boundary longitude and latitude data set of the area to be selected, create a vector map of the area to be selected, and perform grid processing on the vector map of the area to be selected to obtain a grid distribution map; Obtain a sampling address data set in the grid distribution map through an address uniform sampling mechanism; Based on the comprehensive evaluation data set and the sampling address data set, determine a sampling address annotation data set; According to the distribution law of the point position labels of the sampling address points in the sampling address annotation data set, select a corresponding auxiliary site selection model; Train and optimize the auxiliary site selection model through the sampling address annotation data set to obtain a final auxiliary site selection model for making auxiliary site selection decisions through the final auxiliary site selection model.

2. The machine learning-based auxiliary site selection method according to claim 1, characterized in that, Based on the site selection requirements, determine various types of effective facilities within the area to be selected, and construct an evaluation index system for effective facilities, specifically including: Based on the site selection requirements and the usage nature of the facility to be selected, determine effective facilities of several usage types that affect the site selection decision; among them, the effective facilities of the several usage types at least include: residential providing facilities, medical service facilities, educational service facilities, shopping service facilities, and transportation service facilities; According to the relationship between the effective facilities and the facility to be selected, determine the type of influence of the effective facilities on the site selection decision; among them, the relationship between the effective facilities and the facility to be selected includes a dependence relationship and a competition relationship; the type of influence includes a positive influence and a negative influence; According to the type of influence, determine the category of the role of the effective facilities; among them, the category of the role of the effective facilities corresponding to the positive influence is a positive effective facility, and the category of the role of the effective facilities corresponding to the negative influence is a negative effective facility; Based on the facility attribute characteristics of the effective facilities of each usage type, construct an evaluation index system corresponding to the effective facilities of each usage type; among them, the evaluation indicators in the evaluation index system include: relevant indicators of the population living inside the facility, relevant indicators of the population receiving services inside the facility, relevant indicators of the service content of the facility, and relevant indicators of the internal attributes of the facility.

3. The machine learning-based auxiliary site selection method according to claim 1, characterized in that, Normalize the evaluation data of the same type of effective facilities to obtain a comprehensive evaluation data set corresponding to each type of effective facility, specifically including: Obtain the spatial position data and evaluation data of various types of effective facilities within the area to be selected; among them, the spatial position data at least includes the longitude and latitude of the effective facilities; the evaluation data at least includes the evaluation index data of the effective facilities in their evaluation index system. Divide each evaluation index in the evaluation index system into a positive evaluation index and a negative evaluation index; Based on the linear transformation algorithm, perform linear transformation on the evaluation data corresponding to the positive evaluation index and the evaluation data corresponding to the negative evaluation index respectively, and convert the evaluation data into standard evaluation data within the interval [0,1] with the same action direction to obtain a standard evaluation data set; Determine the weights of each evaluation index in the evaluation index system of various effective facilities through the entropy weight method; According to the weights and the standard evaluation data set, calculate the comprehensive evaluation values of various effective facilities through the linear weighting method to obtain a comprehensive evaluation data set of effective facilities; wherein, the comprehensive evaluation data set at least includes: the longitude and latitude coordinates of the effective facilities, the comprehensive evaluation value, and the action category.

4. The auxiliary site selection method based on machine learning according to claim 1, characterized in that, Obtain the boundary longitude and latitude data set of the area to be site-selected, create a vector map of the area to be site-selected, and perform grid processing on the vector map of the area to be site-selected to obtain a grid distribution map, specifically including: Obtain the boundary longitude and latitude data of all administrative regions involved in the area to be site-selected to obtain the boundary longitude and latitude data set of the area to be site-selected; If there are incomplete administrative regions involved in the area to be site-selected, then delimit a polygon polyline area containing the area to be site-selected in the incomplete administrative region, and determine the boundary longitude and latitude data set of the area to be site-selected through the longitude and latitude coordinates of the polygon polyline area; Create a vector map of the area to be site-selected based on the boundary longitude and latitude data set; According to the map information in the vector map of the area to be site-selected, the coverage radius of the facility to be site-selected, and the sampling address point sampling density requirements, set the type, creation range, horizontal interval, and vertical interval of the grid area, and select the same reference coordinate system as the vector map of the area to be site-selected to create a grid for the vector map of the area to be site-selected to obtain the grid distribution map; wherein, the map information in the vector map of the area to be site-selected includes: the range and area of the area to be site-selected, and the association rule between longitude and latitude and distance within the area to be site-selected.

5. The auxiliary site selection method based on machine learning according to claim 1, characterized in that, Obtain a sampling address data set in the grid distribution map through an address uniform sampling mechanism, specifically including: Determine the basic sampling dot matrix corresponding to each grid unit in the grid distribution map according to the address uniform sampling mechanism; the basic sampling dot matrix includes 1 central sampling address point and corresponding 4 associated sampling address points, and a digital mark corresponding to each sampling address point; Calculate the centroid longitude and latitude coordinates of each grid unit according to the longitude and latitude data of the grid side line to obtain the longitude and latitude coordinates of the central sampling address point; Calculate the longitude and latitude coordinates of the associated sampling address points according to the association rule between the central sampling point and the longitude and latitude coordinates of the associated sampling address points; and determine the digital labels of each sampling address point according to the digital label generation rule to obtain the sampling address data set; the sampling address data set includes: the longitude and latitude coordinates and digital labels of each sampling address point.

6. A machine learning-based auxiliary site selection method according to claim 1, characterized in that Based on the comprehensive evaluation data set and the sampling address data set, determine the sampling address annotation data set, specifically including: Based on the comprehensive evaluation data set and the sampling address data set, use the discount quantization function of the comprehensive evaluation value of effective facilities: Calculate the discount coefficient Φ(p,q) corresponding to the comprehensive evaluation value of effective facilities for each sampling address point; where r is the coverage radius of the facility to be located; λ ∈ (0,1) is the proportion of the population served within half of the coverage radius to the total population served; d(p,q) is the distance between the sampling address point p and the associated effective facility q. According to Calculate the forward decision correlation data of the sampling address point where p is the sampling address point and q + is a forward effective facility, and Φ(p, q + ) is the discount coefficient corresponding to the comprehensive evaluation value of the effective facility q + for the sampling address point p, is the comprehensive evaluation value of the effective facility q + ; According to Calculate the negative decision - associated data of the sampling address point where p is the sampling address point and q - is a negative effective facility, and Φ(p, q - ) is the discount coefficient corresponding to the comprehensive evaluation value of the effective facility q - for the sampling address point p, and is the comprehensive evaluation value of the effective facility q - ; Synchronize the obtained positive decision association data and negative decision association data to the sampling address data set to obtain a sampling address decision association data set; the sampling address decision association data set at least includes: the longitude and latitude coordinates, digital labels, positive decision association data and negative decision association data of the sampling address points; Determine the label annotation method for the sampling address points, and use the annotation method to perform label annotation on the sampling address decision association data set to obtain the sampling address annotation data set; wherein, the sampling address annotation data set at least includes: the longitude and latitude coordinates, marked numbers and point position labels of the sampling address points.

7. A machine learning-based auxiliary site selection method according to claim 6, characterized in that Determine the label annotation method for the sampling address points, and use the annotation method to perform label annotation on the sampling address decision association data set to obtain the sampling address annotation data set, specifically including: Determine the label form of the sampling address points according to the site selection decision feedback form; wherein, the site selection decision feedback form includes a scoring feedback form and a rating feedback form; the label form includes a scoring label form and a rating label form; If the determined label form is the scoring label form, calculate the scoring data of each sampling address point in the sampling address decision association data set based on the scoring function, and label the scoring data as the point position label of the sampling address point; If the determined label form is the rating label form, calculate the weighted decision association data of each sampling address point in the sampling address decision association data set based on the weighting function; based on the rating function, map the weighted decision association data to the corresponding rating data, and label the rating data as the point position label of the sampling address point.

8. A machine learning-based auxiliary site selection method according to claim 1, characterized in that According to the distribution rule of the point position labels of the sampling address points in the sampling address annotation data set, select the corresponding auxiliary site selection model, specifically including: Analyze the distribution rule of the point position labels of each sampling address point in the sampling address annotation data set, determine that the point position label distribution rule is a non-linear distribution, and initially select an auxiliary site selection model of the non-linear model type; If the sampling address point label is in the scoring label form, the selected auxiliary site selection model is a non-linear regression model; if the sampling address point label is in the rating label form, the selected auxiliary site selection model is a non-linear classification model.

9. A machine learning-based auxiliary site selection method according to claim 1, It is characterized in that the auxiliary site selection model is trained and optimized by labeling the data set with the sampling address to obtain the final auxiliary site selection model, specifically including: According to the digital marks of each sampling address point in the sampling address labeled data set, determine the splitting scheme of the sampling address labeled data set, and split the sampling address labeled data set into a training data set, a validation data set and a test data set; Through the training data set and the validation data set, perform hyperparameter tuning on the auxiliary site selection model to determine the hyperparameter combination with the optimal performance; Merge the training data set and the validation data set into a new training data set; Based on the new training data set and the hyperparameter combination, train the auxiliary site selection model to obtain the final auxiliary site selection model; Based on the test data set, calculate the evaluation index of the final auxiliary site selection model to test and evaluate the performance of the final auxiliary site selection model.

10. An auxiliary site selection device based on machine learning It is characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, so that the at least one processor can execute an auxiliary site selection method based on machine learning according to any one of claims 1-9.

Citation Information

Patent Citations

  • Self-adaptive site selection evaluation method and device based on BP neural network, and medium

    CN118195683A