Site selection method, device, equipment and storage medium based on point of interest data

By analyzing the competition and collaboration relationships of point of interest data and calculating the attraction between brands and communities, the problem that existing site selection methods fail to effectively handle complex competition-collaboration relationships is solved, and more accurate site selection recommendations are achieved.

CN119884451BActive Publication Date: 2025-09-26INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411955117.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-09-26
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Existing site recommendation methods fail to effectively deal with complex competition-collaboration relationships, ignore the influence between brands, lack flexibility, find it difficult to comprehensively consider the collaborative effect and competitive suppression effect in multi-store site selection, and the calculation of the impact range is not accurate enough.

Method used

By collecting national point of interest data, analyzing business data, resident data and competitive relationships, calculating the attractiveness score and correction coefficient of the points of interest, and comprehensively considering the attraction between the brand and the community, candidate addresses within the region are selected.

Benefits of technology

It improves the adaptability of the site selection model to complex scenarios, realizes the gradual optimization of multi-store site selection problems, avoids the decline in global efficiency caused by the influence of previous site selection, and provides more practical site recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884451B_ABST
    Figure CN119884451B_ABST
Patent Text Reader

Abstract

The present disclosure provides a site selection method, apparatus, device and storage medium based on point of interest data, the method comprising: collecting point of interest data nationwide and obtaining target data, retrieving at least one first point of interest data and corresponding first related data corresponding to the point of interest data based on the target data; calculating an attraction score corresponding to at least one first point of interest data based on business data; setting a first correction coefficient based on other point of interest related data and the first related data; calculating a comprehensive scale weight of the at least one first point of interest data based on the attraction score and the first correction coefficient; retrieving resident data in the first related data based on the location information, calculating the community attraction of the at least one first point of interest data based on the resident data and the comprehensive scale weight, and obtaining a candidate address within the area based on the community attraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data retrieval, and in particular to a site selection method, apparatus, device, and storage medium based on point of interest data. Background Art

[0002] In the task of site recommendation, traditional methods typically only consider single metrics (such as customer flow and rent) and lack a comprehensive analysis of competing and collaborating brand relationships. Existing solutions struggle to effectively model the impact of complex competitive and collaborative relationships on site selection, especially with limited analysis of the overlapping impacts of different brands within the same category. They ignore the impact of competing brands and fail to fully consider the synergistic effects of collaborating brands. They fail to incorporate the competitive and collaborative relationships between brands, resulting in inaccurate calculations of the impact area. They rely solely on inferences based on existing data, unable to handle new brands or the dynamics of competitive relationships, and lack flexibility. They also lack quantitative modeling of the relationships between competing and collaborating brands. They struggle to comprehensively consider the collaborative and competitive suppression effects in multi-store site selection. Their impact area calculations are crude, ignoring important factors such as brand characteristics, category weights, and distance decay. Summary of the Invention

[0003] The present disclosure provides a site selection method, apparatus, device and storage medium based on point of interest data to at least solve the above technical problems existing in the prior art.

[0004] According to a first aspect of the present disclosure, a site selection method based on point of interest data is provided, the method comprising:

[0005] collecting national point of interest data based on web crawler technology, retrieving relevant data based on the point of interest data, and storing the relevant data and the point of interest data in a database;

[0006] Acquire target data, and retrieve corresponding at least one first point of interest data and corresponding first related data from the database based on the target data;

[0007] searching for business data corresponding to the at least one first point of interest data based on the first related data, and calculating an attractiveness score corresponding to the at least one first point of interest data based on the business data;

[0008] obtaining location information of at least one first point of interest data in the first related data, retrieving other point of interest related data in the related data based on the location information, and setting a first correction coefficient based on the other point of interest related data and the first related data;

[0009] Calculating a comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient;

[0010] Based on the location information, resident data in the first relevant data is retrieved, and the community attraction of the at least one first point of interest data is calculated based on the resident data and the comprehensive scale weight. The community attraction of the at least one first point of interest data is accumulated to obtain the overall attraction of the first point of interest data in the area corresponding to the target data, and the candidate addresses in the area are obtained based on the overall attraction and the community attraction.

[0011] In one possible implementation, calculating the attractiveness score corresponding to the at least one first point of interest data based on the business data includes:

[0012] The business data includes: evaluation data, business scale data, and business category data;

[0013] Setting evaluation weights for the number of reviews and the star ratings in the evaluation data, respectively, and calculating a brand influence score corresponding to the at least one first point of interest data based on the evaluation weights, the number of reviews, and the star ratings;

[0014] Based on the business scale data, the brand attraction value corresponding to the at least one first point of interest data is calculated; based on the business category data, at least one category weight is set; the first category weight corresponding to the first point of interest data is selected from the at least one category weight; and based on the first category weight, the brand attraction value and the brand influence score, the attraction score corresponding to the at least one first point of interest data is calculated.

[0015] In one embodiment, the setting of the first correction coefficient based on the other point of interest related data and the first related data includes:

[0016] The data related to other points of interest include actual customer flow data and actual consumption data of customers corresponding to the other points of interest;

[0017] Setting expected passenger flow data and expected consumption data corresponding to the first point of interest based on the first relevant data;

[0018] Calculating the competition intensity of the point of interest based on the actual passenger flow data, the actual consumption data, the expected passenger flow data, and the expected consumption data;

[0019] The distance between points of interest is obtained based on the first relevant data and the other point of interest related data, a competition distance attenuation coefficient is set based on the distance between points of interest, and the first correction coefficient is set based on the competition distance attenuation coefficient, the distance between points of interest and the point of interest competition intensity.

[0020] In one embodiment, the calculating the community attractiveness of the at least one first point of interest data based on the resident data and the comprehensive scale weight includes:

[0021] The resident data includes: resident portrait feature data and resident traffic data;

[0022] constructing a target group feature vector based on the at least one first point of interest data and the first related data, extracting a resident feature vector based on the resident portrait feature, and calculating a similarity between the resident feature vector and the target group feature vector;

[0023] Classifying the resident traffic data based on traffic categories, setting corresponding traffic mode weights based on the classification results, and calculating weighted reachable distances based on the traffic distances corresponding to the classification results and the traffic mode weights;

[0024] The community attractiveness of the at least one first point of interest data is calculated based on the similarity, the weighted reachable distance, and the comprehensive scale weight.

[0025] In one possible implementation manner, obtaining candidate addresses within the area based on the overall attractiveness and the community attractiveness includes:

[0026] Obtain a set of optional addresses within the area, and use each of the optional addresses as a seed node;

[0027] Obtaining a total attraction set corresponding to all points of interest in the target data, sorting the values ​​in the total attraction set, and selecting a seed node corresponding to the point of interest with the largest value as a single candidate address;

[0028] An adjustment coefficient is set according to the type of the point of interest. Based on the adjustment coefficient, the overall attraction set and the competition intensity of the points of interest, the attraction of the points of interest after removing the point of interest with the largest value is calculated, and the attraction of the remaining points of interest is repeatedly calculated. After the number of repeated iterations reaches the set value, the nodes corresponding to the calculation results are selected as multiple candidate addresses.

[0029] According to a second aspect of the present disclosure, a site selection device based on point of interest data is provided, the device comprising:

[0030] a data collection unit, configured to collect national point of interest data based on a web crawler technology, retrieve relevant data based on the point of interest data, and store the relevant data and the point of interest data in a database;

[0031] a target data retrieval unit, configured to obtain target data, and retrieve corresponding at least one first point of interest data and corresponding first related data from the database based on the target data;

[0032] an attraction score calculation unit, configured to search for business data corresponding to the at least one first point of interest data based on the first related data, and calculate an attraction score corresponding to the at least one first point of interest data based on the business data;

[0033] a first correction coefficient calculation unit, configured to obtain location information of at least one first point of interest data in the first related data, retrieve other point of interest related data in the related data based on the location information, and set a first correction coefficient based on the other point of interest related data and the first related data;

[0034] a comprehensive scale weight calculation unit, configured to calculate a comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient;

[0035] an attraction calculation unit, configured to retrieve resident data from the first relevant data based on the location information, calculate the community attraction of the at least one first point of interest data based on the resident data and the comprehensive scale weight, and accumulate the community attraction of the at least one first point of interest data to obtain an overall attraction of the first point of interest data within the area corresponding to the target data;

[0036] A site selection unit is configured to obtain candidate addresses within the area based on the overall attractiveness and the community attractiveness.

[0037] In one possible implementation, the attraction score calculation unit is further configured to: the business data includes: evaluation data, business scale data, and business category data; set evaluation weights for the number of reviews and rating stars in the evaluation data, respectively; calculate the brand influence score corresponding to the at least one first point of interest data based on the evaluation weights, the number of reviews, and the rating stars; calculate the brand attraction value corresponding to the at least one first point of interest data based on the business scale data, set at least one category weight based on the business category data, select the first category weight corresponding to the first point of interest data from the at least one category weight, and calculate the attraction score corresponding to the at least one first point of interest data based on the first category weight, the brand attraction value, and the brand influence score;

[0038] The attraction calculation unit is also used for: the resident data includes: resident portrait feature data and resident traffic data; constructing a target group feature vector based on the at least one first point of interest data and the first related data, extracting a resident feature vector based on the resident portrait feature, and calculating the similarity between the resident feature vector and the target group feature vector; classifying the resident traffic data based on traffic categories, setting corresponding traffic mode weights based on the classification results, and calculating the weighted reachable distance based on the traffic distance corresponding to the classification results and the traffic mode weight; calculating the community attraction of the at least one first point of interest data based on the similarity, the weighted reachable distance and the comprehensive scale weight.

[0039] In one embodiment, the first correction coefficient calculation unit is further configured to: the other point-of-interest related data include actual passenger flow data and actual consumption data of customers corresponding to the other point-of-interest; set expected passenger flow data and expected consumption data corresponding to the first point of interest based on the first related data; calculate the point-of-interest competition intensity based on the actual passenger flow data, the actual consumption data, the expected passenger flow data, and the expected consumption data; obtain the distance between points of interest based on the first related data and the other point-of-interest related data, set a competition distance attenuation coefficient based on the distance between points of interest, and set the first correction coefficient based on the competition distance attenuation coefficient, the distance between points of interest, and the point-of-interest competition intensity;

[0040] The site selection unit is also used to obtain a set of optional addresses in the area, and take each of the optional addresses as a seed node; obtain a set of overall attractions corresponding to all the points of interest in the target data, sort the values ​​in the overall attraction set, and select the seed node corresponding to the point of interest with the largest value as a single candidate address; set an adjustment coefficient according to the type of the point of interest, and calculate the attraction of the point of interest after removing the point of interest with the largest value based on the adjustment coefficient, the overall attraction set and the competition intensity of the point of interest, and repeatedly iterate the calculation of the attraction of the remaining points of interest. After the number of repeated iterations reaches the set value, select the nodes corresponding to the calculation results as multiple candidate addresses.

[0041] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0042] at least one processor; and

[0043] a memory communicatively connected to the at least one processor; wherein,

[0044] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.

[0045] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.

[0046] This disclosure discloses a site selection method, device, equipment, and storage medium based on point-of-interest data. By analyzing the competitive and collaborative relationships between different brands in target data, the method calculates the attractiveness between brands and communities, and selects candidate addresses within a region based on the calculated results. This integration of competitive and collaborative relationships improves the site selection model's adaptability to complex scenarios, enabling gradual optimization of multi-store site selection, avoiding the global efficiency drop caused by ignoring the impact of previous site selections, and achieving more practical recommendations by incorporating rental costs and customer flow into site recommendations.

[0047] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:

[0049] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.

[0050] Figure 1 The following is a schematic diagram showing the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure. Figure 1 ;

[0051] Figure 2 The following is a schematic diagram showing the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure. Figure 2 ;

[0052] Figure 3 A schematic diagram of a site selection device based on point of interest data according to an embodiment of the present disclosure is shown;

[0053] Figure 4 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0054] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.

[0055] Figure 1 The following is a schematic diagram showing the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure. Figure 1 ,like Figure 1 As shown, the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure includes the following steps:

[0056] Step 101 : collecting national point of interest data based on web crawler technology, retrieving related data based on the point of interest data, and storing the related data and the point of interest data in a database.

[0057] In the disclosed embodiment, national point of interest (POI) data is collected through web crawler technology, commercial databases, social media and other channels. The relevant data includes: relevant business data, including latitude and longitude, address, province, city, district, industry, business district, region, and population information.

[0058] In the disclosed embodiment, the collected data is cleaned, dirty data is removed, and the data is unified in format and integrated before being classified and stored in a database to ensure data quality and facilitate subsequent analysis and query.

[0059] Step 102: Acquire target data, and retrieve corresponding at least one first point of interest data and corresponding first related data from the database based on the target data.

[0060] In the embodiment of the present disclosure, the target data is the target market and brand, and at least one point of interest data is POI data related to the target brand or market. Correspondingly, the first related data is the related business data involved in the POI data related to the target brand or market, including latitude and longitude, address, province, city, district, industry, business district, region, population information, etc.

[0061] Step 103: searching for business data corresponding to the at least one first point of interest data based on the first related data, and calculating an attraction score corresponding to the at least one first point of interest data based on the business data.

[0062] In an embodiment of the present disclosure, business data includes: evaluation data, business scale data, and business category data, wherein the evaluation data is divided into star ratings and number of reviews, as well as market rankings, brand research reports, etc.; evaluation weights are set for the number of reviews and the star ratings in the evaluation data according to the number of reviews and the level of the star ratings, and the brand influence score corresponding to the at least one first point of interest data is calculated based on the evaluation weights, the number of reviews, and the star ratings; business scale data includes business area, shopping mall area, number of parking spaces, etc., and the brand attraction value corresponding to the at least one first point of interest data is calculated based on the business scale data, and at least one category weight is set based on the business category data, wherein business categories can be divided into: large shopping malls, supermarkets, restaurants, etc., and the first category weight corresponding to the first point of interest data is selected from the at least one category weight, and the attraction score corresponding to the at least one first point of interest data is calculated based on the first category weight, the brand attraction value, and the brand influence score.

[0063] Step 104: Obtain location information of the at least one first point of interest data in the first related data, retrieve other point of interest-related data in the related data based on the location information, set a first correction coefficient based on the other point of interest-related data and the first related data, and calculate a comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient.

[0064] In the embodiment of the present disclosure, other points of interest related data are points of interest data of other brands within the same location range, which can be competing products or unrelated brands. Other points of interest related data include actual passenger flow data and actual consumption data of customers corresponding to the other points of interest; based on the first related data, the expected passenger flow data and expected consumption data corresponding to the first point of interest are set; the competition intensity of the points of interest is calculated based on the actual passenger flow data, the actual consumption data, the expected passenger flow data and the expected consumption data, and the relationship between the points of interest can be determined according to the magnitude of the competition intensity of the points of interest, for example, if the competition intensity of the points of interest is greater than zero, it indicates a complementary relationship, and if the competition intensity of the points of interest is less than zero, it indicates a competitive relationship. The distance between the points of interest is obtained based on the first related data and the other points of interest related data, wherein the distance between the points of interest can be a physical distance or a time distance, and a competition distance attenuation coefficient is set based on the distance between the points of interest to control the degree of weakening of the intensity of collaboration or competition with distance. Considering the influence of competition-collaboration, an adjustment coefficient can also be set based on experience to control the amplitude of the overall collaboration-competition correction, and the first correction coefficient is set based on the competition distance attenuation coefficient, the distance between the points of interest, the adjustment coefficient and the competition intensity of the points of interest.

[0065] In the embodiment of the present disclosure, the comprehensive scale weight of the at least one first point of interest data is calculated based on the attractiveness score and the first correction coefficient, wherein the brand basic attractiveness (such as brand influence, business scale, see above for details); the first correction coefficient is used to measure the collaborative and competitive influence of other POIs around the POI.

[0066] Step 105: retrieve the resident data in the first relevant data based on the location information, calculate the community attraction of the at least one first point of interest data based on the resident data and the comprehensive scale weight, accumulate the community attraction of the at least one first point of interest data to obtain the overall attraction of the first point of interest data in the area corresponding to the target data, and obtain the candidate addresses in the area based on the overall attraction and the community attraction.

[0067] In the disclosed embodiment, the resident data includes: resident portrait feature data and resident traffic data; in order to analyze the matching degree between the community resident portrait (age, occupation, income, etc.) and the POI service target group, a target group feature vector is constructed based on the at least one first point of interest data and the first related data, a resident feature vector is extracted based on the resident portrait feature, and the similarity between the resident feature vector and the target group feature vector is calculated; in order to comprehensively consider the accessibility of walking, private car and public transportation, the resident traffic data is classified based on the traffic category, and the corresponding traffic mode weights are set based on the classification results, and the weighted accessible distance is calculated based on the traffic distance corresponding to the classification result and the traffic mode weight; according to the fact that attraction is proportional to resource scale and inversely proportional to distance, the community attractiveness of the at least one first point of interest data is calculated based on the similarity, the weighted accessible distance and the comprehensive scale weight. The community attractiveness of the at least one first point of interest data is accumulated to obtain the overall attractiveness of the first point of interest data in the area corresponding to the target data.

[0068] In an embodiment of the present disclosure, a set of optional addresses in the area is obtained, and each of the optional addresses is used as a seed node; the overall attractiveness of all points of interest in the area is calculated using the above method to obtain a set of overall attractiveness corresponding to all points of interest in the target data, the values ​​in the overall attractiveness set are sorted, and the seed node corresponding to the point of interest with the largest value is selected as a single candidate address; if multiple candidate addresses need to be selected, an adjustment coefficient is set according to the type of the point of interest, wherein the type of the point of interest can be: catering industry, retail industry, logistics industry, etc. Based on the adjustment coefficient, the overall attractiveness set and the competition intensity of the points of interest, the attractiveness of the points of interest after removing the point of interest with the largest value is calculated, and the attractiveness of the remaining points of interest is repeatedly iterated. After the number of repeated iterations reaches a set value, wherein the number of iterations is the same as the number of candidate addresses, the nodes corresponding to the calculation results are selected as multiple candidate addresses.

[0069] Figure 2 The following is a schematic diagram showing the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure. Figure 2 ,like Figure 2 As shown, the implementation process of a site selection method based on point of interest data in an embodiment of the present disclosure includes the following steps:

[0070] Step 201: data import.

[0071] In this disclosed embodiment, national POI data and related commercial data, including latitude and longitude, address, province, city, district, industry, business district, region, and population information, are collected through web crawler technology, commercial databases, social media, and other channels. The data is then cleaned to remove dirty data, unified into a unified format, and stored in a database after classification to ensure data quality and facilitate subsequent analysis and query.

[0072] Step 202: Calculate the candidate point impact gain.

[0073] In the embodiment of the present disclosure, firstly, the basic attraction (B i ) calculation, basic attractiveness mainly reflects the inherent characteristics of the brand and location, and is usually a static value related to brand influence, business scale and business type. Calculating basic attractiveness requires calculating brand influence, business scale and business type separately. Specifically:

[0074] Calculate brand influence (B brand,i ): Set a rating value based on brand awareness and market evaluation. Data sources include: user reviews (such as star ratings, number of reviews), market rankings, brand research reports, etc. The example formula is:

[0075] B brand,i=w1·StarRating i +w2·log(ReviewCount i )

[0076] Among them: StarRating i : The brand's average star rating (e.g. 4.5 stars). ReviewCount i : Number of reviews (e.g. 1000). w1, w2: Weights, used to balance the impact of star rating and number of reviews.

[0077] Calculate the business scale (B size,i ) : Assigns an attractiveness value based on the business area or facility size of the POI. Data sources include: mall area, number of seats, number of parking spaces, etc. The example formula is:

[0078] B size,i =log(Area i )

[0079] Where: Area i is the business area of ​​the POI (square meters), and logarithmic processing is used to reduce the impact of large values.

[0080] Calculation business type (B type,i ): Assign corresponding business type values ​​based on the brand's business category and industry standards or experience values. For example, the basic appeal of large supermarkets and comprehensive shopping malls may be higher than that of small stores with a single business format. For example: Large shopping mall: B type,i =1.2. Supermarket: B type,i =1. Restaurant: B type,i =0.8.

[0081] Calculate the basic attraction, the formula is: B i =B brand,i +B size,i +B type,i .

[0082] In the embodiment of the present disclosure, the comprehensive scale weight (S) is calculated based on the above basic attractiveness and the collaboration-competition correction coefficient. i ), where the collaboration-competition correction factor (C i ) is calculated as follows:

[0083]

[0084] Among them, POIs of different brands i and POI k The intensity of collaboration and competition is defined by the extended Customer Satisfaction Index (CSI), where CSI ik>0: indicates collaborative relationship (complementary effect); CSI ik <0: indicates a competitive relationship (substitution effect). ik :POI i and POI k The physical or temporal distance between the two indicates the degree to which the intensity of cooperation-competition decays with distance. γ: distance decay coefficient, which controls the degree to which the intensity of cooperation or competition decays with distance. A larger γ indicates that the influence of cooperation-competition at long distances is rapidly weakened. δ: adjustment coefficient, which is used to control the magnitude of the overall cooperation-competition correction: δ>0: means that the influence of cooperation and competition is included in the attraction correction. δ=0: means that the cooperation-competition relationship is ignored. According to POI i The collaboration-competition correction factor (C i ) It can be seen that C i >1: indicates that the overall collaborative effect is dominant; C i <1: indicates that the overall competition effect is dominant. ik >0 and d ik When it is small, POI k POI i The collaborative impact is greater, and the correction factor C i Increase, attractiveness improves. When CSI ik <0 and d ik When it is small, POI k POI i The competition influence is greater, and the correction factor C i The influence of cooperation and competition decreases with the distance d. ik Increase and gradually weaken, If CSI ik ≈0, indicating that there is no obvious cooperation or competition between the two, and it does not affect C i .

[0085] In the embodiment of the present disclosure, the attraction calculation is realized based on the gravity model, and the attraction is proportional to the resource scale and inversely proportional to the distance. Set β as the distance attenuation coefficient (experience value, usually between 1.0 and 2.0), POI i The attractiveness G to a community j ij Expressed as:

[0086]

[0087] Select POI i The attraction gain is:

[0088]

[0089] In the embodiment of the present disclosure, R ijTo measure the degree of demand matching between POI and the target community group, the specific calculation method is as follows: Demand matching measures whether the POI's services meet the needs of the target community residents, usually based on community population, income, consumption habits and other characteristics. By analyzing the matching degree between the community resident profile (age, occupation, income, etc.) and the POI service target group, the resident feature vector is obtained, and the target group feature vector is set according to the target POI (such as suitable for children / young people / elderly people). The formula is:

[0090] R ij =Sim(Profile i ,Target i )

[0091] Among them, Profile i Target: The characteristic vector of residents in community j. i :POI i The target group feature vector. Sim: Similarity function (such as cosine similarity).

[0092] In the embodiment of the present disclosure, d ij To weight the accessible distance, the total distance is defined as follows by comprehensively considering the accessibility of walking, private car and public transportation:

[0093]

[0094] in, The distances of the three modes of transportation. w1, w2, w3: The weight of each mode of transportation (determined by the travel habits of community residents).

[0095] Step 203: sorting candidate points.

[0096] In the embodiment of the present disclosure, all candidate POIs: P = {p1, p2, ..., p n} is regarded as a node, and based on the method in step 202 above, the initial attraction gain G of each node in the set is calculated. i The candidate points are sorted according to the size of the calculation results, and the points with the highest attractiveness, namely the TOP K points, are selected as the priority recommended site selection points.

[0097] Step 204: candidate point selection.

[0098] In the embodiment of the present disclosure, considering the comprehensive influence of potential customer coverage, collaborative and competitive POIs, and the decreasing effect of the selected nodes on the candidate nodes, the result of the above step 203: the highest ranked value is added to the selected set S, and then the gravitational gain G of the remaining candidate points is updated. i, deduct the competitive inhibition effect of the selected nodes from the CSI, and repeat the above steps until the number of sites or coverage targets are met, in which the gravitational gain G of the remaining candidate points is updated i The formula is:

[0099]

[0100] Where, S: selected POI set, δ F : Adjustment coefficient, controlling the overall magnitude of competition and collaboration: δ F >0: Collaborative effect dominates. F <0: Competition effect dominates. F =0: Completely ignore the relationship between cooperation and competition, and the model only considers the basic attractiveness of the candidate point. In actual use, it can be dynamically adjusted: δ F This value can be set based on business needs and scenarios. For example, in the catering industry, δ>0.5 emphasizes the importance of collaborative effects. In the retail industry, δ∈[-0.5,0] balances collaboration and competition, favoring competitive suppression. In the logistics industry, δ≈0 ignores competitive effects and only considers coverage.

[0101] In this embodiment, based on the algorithm in step 204, a practical example is provided to further illustrate the following scenario: a fast-moving consumer goods brand wants to select three new store locations within a city. The candidate POI set P = {A, B, C, D, E}, as well as the potential customer coverage, collaborative POI, and competitive POI distribution of each POI are input. The initial attractiveness gain G is calculated. i :

[0102] G A =120,G B =110,G C =150,G D =140,G E =100

[0103] In the first iteration, select C (highest attraction gain), update the attraction of other points, considering G C The CSI impact of the updated gravitational gain of the remaining candidate points is:

[0104] G′ A =G A -δCSI AC

[0105] G′ B =G B -δCSI BC

[0106] In the second iteration, the point with the highest attractiveness after the update is selected, and the final output recommended site points are C, A, and D.

[0107] Step 205: Visualize and output candidate points.

[0108] In the disclosed embodiments, visualization output can be performed in the following ways, for example: map display: marking POI data on a map to intuitively display the distribution of POIs; heat map display: reflecting the density of POIs in a certain area through the depth of color to help identify hot spots; category display: using different colors or shapes to mark POIs according to their different categories (such as catering, entertainment, shopping, etc.) on the map to facilitate identification of different types of POI distribution.

[0109] In the disclosed embodiments, the hardware architecture of the present invention utilizes a distributed architecture for deployment. Specifically, tasks such as data collection, density calculation, brand analysis, and white space market assessment are distributed to different servers for processing. To ensure processing speed and stability, load balancing technology can be used to manage and schedule the servers. Furthermore, to facilitate enterprise user experience, the algorithm can be integrated into the enterprise's internal business systems, or a standalone web application or mobile application can be developed for user use.

[0110] Figure 3 A schematic diagram of a site selection device based on point of interest data according to an embodiment of the present disclosure is shown. Figure 3 As shown, an embodiment of the present disclosure provides a site selection device based on point of interest data, including:

[0111] The data collection unit 301 is configured to collect national point of interest data based on a web crawler technology, retrieve related data based on the point of interest data, and store the related data and the point of interest data in a database.

[0112] The target data retrieval unit 302 is configured to obtain target data, and retrieve corresponding at least one first point of interest data and corresponding first related data in the database based on the target data.

[0113] The attraction score calculation unit 303 is configured to search for business data corresponding to the at least one first point of interest data based on the first related data, and calculate an attraction score corresponding to the at least one first point of interest data based on the business data.

[0114] The attraction score calculation unit 303 is also used for the business data including: evaluation data, business scale data, and business category data; setting evaluation weights for the number of comments and rating stars in the evaluation data respectively, and calculating the brand influence score corresponding to the at least one first point of interest data based on the evaluation weights, the number of comments, and the rating stars; calculating the brand attraction value corresponding to the at least one first point of interest data based on the business scale data, setting at least one category weight based on the business category data, selecting the first category weight corresponding to the first point of interest data from the at least one category weight, and calculating the attraction score corresponding to the at least one first point of interest data based on the first category weight, the brand attraction value, and the brand influence score.

[0115] The first correction coefficient calculation unit 304 is used to obtain the position information of at least one first point of interest data in the first relevant data, retrieve other point of interest related data in the relevant data based on the position information, and set a first correction coefficient based on the other point of interest related data and the first relevant data.

[0116] The first correction coefficient calculation unit 304 is further configured to: the other POI-related data includes actual customer flow data and actual consumption data of customers corresponding to the other POIs; set expected customer flow data and expected consumption data corresponding to the first POI based on the first related data; calculate the POI competition intensity based on the actual customer flow data, actual consumption data, expected customer flow data, and expected consumption data; obtain the distance between POIs based on the first related data and the other POI-related data; set a competition distance decay coefficient based on the distance between POIs; and set the first correction coefficient based on the competition distance decay coefficient, the distance between POIs, and the POI competition intensity.

[0117] The comprehensive scale weight calculation unit 305 is configured to calculate the comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient.

[0118] The attraction calculation unit 306 is used to retrieve the resident data in the first related data based on the location information, calculate the community attraction of the at least one first point of interest data based on the resident data and the comprehensive scale weight, accumulate the community attraction of the at least one first point of interest data, and obtain the overall attraction of the first point of interest data in the area corresponding to the target data.

[0119] The attraction calculation unit 306 is also used for, wherein the resident data includes: resident portrait feature data and resident traffic data; constructing a target group feature vector based on the at least one first point of interest data and the first related data, extracting a resident feature vector based on the resident portrait feature, and calculating the similarity between the resident feature vector and the target group feature vector; classifying the resident traffic data based on traffic categories, setting corresponding traffic mode weights based on the classification results, and calculating the weighted reachable distance based on the traffic distance corresponding to the classification results and the traffic mode weight; calculating the community attraction of the at least one first point of interest data based on the similarity, the weighted reachable distance and the comprehensive scale weight.

[0120] The site selection unit 307 is configured to obtain candidate addresses within the area based on the overall attractiveness and the community attractiveness.

[0121] The site selection unit 307 is also used to obtain a set of optional addresses in the area, and take each of the optional addresses as a seed node; obtain a set of overall attractions corresponding to all the points of interest in the target data, sort the values ​​in the overall attraction set, and select the seed node corresponding to the point of interest with the largest value as a single candidate address; set an adjustment coefficient according to the type of the point of interest, and calculate the attraction of the point of interest after removing the point of interest with the largest value based on the adjustment coefficient, the overall attraction set and the competition intensity of the point of interest, and repeatedly iterate the calculation of the attraction of the remaining points of interest. After the number of repeated iterations reaches the set value, select the nodes corresponding to the calculation results as multiple candidate addresses.

[0122] In an exemplary embodiment, the data collection unit 301, the target data retrieval unit 302, the attractiveness score calculation unit 303, the first correction coefficient calculation unit 304, the comprehensive scale weight calculation unit 305, the attractiveness calculation unit 306 and the site selection unit 307 can be implemented by one or more central processing units (CPU), graphics processing units (GPU), application specific integrated circuits (ASIC), DSP, programmable logic devices (PLD), complex programmable logic devices (CPLD), field programmable gate arrays (FPGA), general processors, controllers, microcontrollers (MCU), microprocessors, or other electronic components.

[0123] Regarding the device in the above embodiment, the specific manner in which each module and unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0124] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0125] Figure 4 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0126] like Figure 4As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0127] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0128] The computing unit 801 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a site selection method based on point of interest data. For example, in some embodiments, a site selection method based on point of interest data can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the site selection method based on point of interest data described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other appropriate manner (eg, by means of firmware) to execute a location selection method based on point of interest data.

[0129] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0134] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0135] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0136] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0137] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A site selection method based on point of interest data, characterized in that: The method comprises: collecting national point of interest data based on web crawler technology, retrieving relevant data based on the point of interest data, and storing the relevant data and the point of interest data in a database; Acquire target data, and retrieve corresponding at least one first point of interest data and corresponding first related data from the database based on the target data; searching for business data corresponding to the at least one first point of interest data based on the first related data, and calculating an attractiveness score corresponding to the at least one first point of interest data based on the business data; obtaining location information of at least one first point of interest data in the first related data, retrieving other point of interest related data in the related data based on the location information, and setting a first correction coefficient based on the other point of interest related data and the first related data; Calculating a comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient; Based on the location information, resident data in the first relevant data is retrieved, and the community attraction of the at least one first point of interest data is calculated based on the resident data and the comprehensive scale weight. The community attraction of the at least one first point of interest data is accumulated to obtain the overall attraction of the first point of interest data in the area corresponding to the target data, and the candidate addresses in the area are obtained based on the overall attraction and the community attraction.

2. The method according to claim 1, characterized in that The calculating, based on the business data, an attractiveness score corresponding to the at least one first point of interest data includes: The business data includes: evaluation data, business scale data, and business category data; Setting evaluation weights for the number of reviews and the star ratings in the evaluation data, respectively, and calculating a brand influence score corresponding to the at least one first point of interest data based on the evaluation weights, the number of reviews, and the star ratings; Based on the business scale data, the brand attraction value corresponding to the at least one first point of interest data is calculated; based on the business category data, at least one category weight is set; the first category weight corresponding to the first point of interest data is selected from the at least one category weight; and based on the first category weight, the brand attraction value and the brand influence score, the attraction score corresponding to the at least one first point of interest data is calculated.

3. The method according to claim 1, characterized in that The setting of the first correction coefficient based on the other point of interest related data and the first related data includes: The data related to other points of interest include actual customer flow data and actual consumption data of customers corresponding to the other points of interest; Setting expected passenger flow data and expected consumption data corresponding to the first point of interest based on the first relevant data; Calculating the competition intensity of the point of interest based on the actual passenger flow data, the actual consumption data, the expected passenger flow data, and the expected consumption data; The distance between points of interest is obtained based on the first relevant data and the other point of interest related data, a competition distance attenuation coefficient is set based on the distance between points of interest, an adjustment coefficient is set based on the point of interest related data and the first point of interest data, and the first correction coefficient is set based on the competition distance attenuation coefficient, the distance between points of interest, the adjustment coefficient and the point of interest competition intensity.

4. The method according to claim 1, wherein The calculating the community attractiveness of the at least one first point of interest data based on the resident data and the comprehensive scale weight includes: The resident data includes: resident portrait feature data and resident traffic data; constructing a target group feature vector based on the at least one first point of interest data and the first related data, extracting a resident feature vector based on the resident portrait feature, and calculating a similarity between the resident feature vector and the target group feature vector; Classifying the resident traffic data based on traffic categories, setting corresponding traffic mode weights based on the classification results, and calculating weighted reachable distances based on the traffic distances corresponding to the classification results and the traffic mode weights; The community attractiveness of the at least one first point of interest data is calculated based on the similarity, the weighted reachable distance, and the comprehensive scale weight.

5. The method according to claim 3, characterized in that The obtaining of candidate addresses within the area based on the overall attractiveness and the community attractiveness includes: Obtain a set of optional addresses within the area, and use each of the optional addresses as a seed node; Obtaining a total attraction set corresponding to all points of interest in the target data, sorting the values ​​in the total attraction set, and selecting a seed node corresponding to the point of interest with the largest value as a single candidate address; An adjustment coefficient is set according to the type of the point of interest. Based on the adjustment coefficient, the overall attraction set and the competition intensity of the points of interest, the attraction of the points of interest after removing the point of interest with the largest value is calculated, and the attraction of the remaining points of interest is repeatedly calculated. After the number of repeated iterations reaches the set value, the nodes corresponding to the calculation results are selected as multiple candidate addresses.

6. A site selection device based on point of interest data, characterized in that: The device comprises: a data collection unit, configured to collect national point of interest data based on a web crawler technology, retrieve relevant data based on the point of interest data, and store the relevant data and the point of interest data in a database; a target data retrieval unit, configured to obtain target data, and retrieve corresponding at least one first point of interest data and corresponding first related data from the database based on the target data; an attraction score calculation unit, configured to search for business data corresponding to the at least one first point of interest data based on the first related data, and calculate an attraction score corresponding to the at least one first point of interest data based on the business data; a first correction coefficient calculation unit, configured to obtain location information of at least one first point of interest data in the first related data, retrieve other point of interest related data in the related data based on the location information, and set a first correction coefficient based on the other point of interest related data and the first related data; a comprehensive scale weight calculation unit, configured to calculate a comprehensive scale weight of the at least one first point of interest data based on the attractiveness score and the first correction coefficient; an attraction calculation unit, configured to retrieve resident data from the first relevant data based on the location information, calculate the community attraction of the at least one first point of interest data based on the resident data and the comprehensive scale weight, and accumulate the community attraction of the at least one first point of interest data to obtain an overall attraction of the first point of interest data within the area corresponding to the target data; A site selection unit is configured to obtain candidate addresses within the area based on the overall attractiveness and the community attractiveness.

7. The device according to claim 6, characterized in that The attraction score calculation unit is further configured to: the business data includes: evaluation data, business scale data, and business category data; set evaluation weights for the number of reviews and the rating star in the evaluation data, respectively, and calculate the brand influence score corresponding to the at least one first point of interest data based on the evaluation weights, the number of reviews, and the rating star; calculate the brand attraction value corresponding to the at least one first point of interest data based on the business scale data, set at least one category weight based on the business category data, select the first category weight corresponding to the first point of interest data from the at least one category weight, and calculate the attraction score corresponding to the at least one first point of interest data based on the first category weight, the brand attraction value, and the brand influence score; The attraction calculation unit is also used for: the resident data includes: resident portrait feature data and resident traffic data; constructing a target group feature vector based on the at least one first point of interest data and the first related data, extracting a resident feature vector based on the resident portrait feature, and calculating the similarity between the resident feature vector and the target group feature vector; classifying the resident traffic data based on traffic categories, setting corresponding traffic mode weights based on the classification results, and calculating the weighted reachable distance based on the traffic distance corresponding to the classification results and the traffic mode weight; calculating the community attraction of the at least one first point of interest data based on the similarity, the weighted reachable distance and the comprehensive scale weight.

8. The device according to claim 6, characterized in that The first correction coefficient calculation unit is further configured to: the other point-of-interest related data include actual passenger flow data and actual consumption data of customers corresponding to the other points of interest; set expected passenger flow data and expected consumption data corresponding to the first point of interest based on the first related data; calculate the competition intensity of the point of interest based on the actual passenger flow data, the actual consumption data, the expected passenger flow data, and the expected consumption data; obtain the distance between points of interest based on the first related data and the other point-of-interest related data, set a competition distance attenuation coefficient based on the distance between points of interest, set an adjustment coefficient based on the point-of-interest related data and the first point-of-interest data, and set the first correction coefficient based on the competition distance attenuation coefficient, the distance between points of interest, the adjustment coefficient, and the competition intensity of the point of interest; The site selection unit is also used to obtain a set of optional addresses in the area, and take each of the optional addresses as a seed node; obtain a set of overall attractions corresponding to all the points of interest in the target data, sort the values ​​in the overall attraction set, and select the seed node corresponding to the point of interest with the largest value as a single candidate address; set an adjustment coefficient according to the type of the point of interest, and calculate the attraction of the point of interest after removing the point of interest with the largest value based on the adjustment coefficient, the overall attraction set and the competition intensity of the point of interest, and repeatedly iterate the calculation of the attraction of the remaining points of interest. After the number of repeated iterations reaches the set value, select the nodes corresponding to the calculation results as multiple candidate addresses.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Passenger flow data optimization method and system based on interest point filtering and computer readable storage medium

    CN116258250A

  • Interest point processing method and device, electronic equipment and storage medium

    CN117493639A