Merchant location data processing methods, devices, equipment and storage media

By clustering and comprehensively scoring merchant location data from multiple channels, the problem of inaccurate positioning caused by inconsistent channel data quality was solved, achieving higher accuracy and precision retention in merchant positioning.

CN117033731BActive Publication Date: 2025-10-31CHINA UNIONPAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311001018.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-09
Publication Date
2025-10-31
Estimated Expiration
2043-08-09

AI Technical Summary

Technical Problem

The quality of merchant location data from different sources varies, which reduces the accuracy of merchant positioning and makes it difficult to determine the accurate location when using data from multiple sources.

Method used

By clustering the location data sets of each merchant's channel, calculating its own quality score and cross-referenced quality scores, and comprehensively considering the overall quality score of the data from each channel, the merchant's location is determined using the data with the highest overall quality.

Benefits of technology

It improves the accuracy of merchant positioning, reduces the adverse effects of inconsistent and conflicting data quality from different channels on positioning, preserves the accuracy of location data, and avoids accuracy loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117033731B_ABST
    Figure CN117033731B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for processing merchant location data, relating to the field of data processing. The method includes: clustering location data sets from each channel source for a target merchant to obtain clusters for each location data set, where each location data set includes location data; obtaining a quality score for each location data set based on the clusters and preset data quality rules; obtaining a cross-reference quality score for each location data set relative to other location data sets based on the clusters; obtaining a comprehensive quality score for the location data sets based on their own quality scores and cross-reference quality scores; and determining the location of the target merchant based on the location data set with the highest quality, as represented by the comprehensive quality score. Embodiments of this application can improve the accuracy of merchant location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing, and in particular relates to a method, apparatus, device and storage medium for processing merchant location data. Background Technology

[0002] With the continuous development of information technology, more and more merchants' information services need to be based on their location. For merchants, location data can be obtained from multiple sources, and the accurate location of the merchant can be determined based on the location data from multiple sources. However, the quality of location data from different sources varies, and there may be conflicts between them, making it difficult to obtain the accurate location of the merchant even when location data from multiple sources is available, thus reducing the accuracy of merchant location positioning. Summary of the Invention

[0003] This application provides a merchant location data processing method, apparatus, device, and storage medium that can improve the accuracy of merchant location.

[0004] In a first aspect, embodiments of this application provide a merchant location data processing method, comprising: clustering location data sets from each channel source of a target merchant to obtain clusters for each location data set, wherein the location data set includes location data; obtaining a quality score for each location data set based on the clusters and preset data quality rules; obtaining a cross-reference quality score for each location data set relative to other location data sets based on the clusters; obtaining a comprehensive quality score for the location data sets based on their own quality scores and cross-reference quality scores; and determining the location of the target merchant based on the location data set with the highest quality as represented by the comprehensive quality score.

[0005] Secondly, embodiments of this application provide a merchant location data processing apparatus, comprising: a clustering module, configured to cluster location data sets from each channel source of a target merchant to obtain a cluster of each location data set, wherein the location data set includes location data; a first score determination module, configured to obtain a self-quality score for each location data set based on the cluster and preset data quality rules; a second score determination module, configured to obtain a cross-reference quality score of each location data set relative to other location data sets based on the cluster; a comprehensive score determination module, configured to obtain a comprehensive quality score for the location data set based on its self-quality score and cross-reference quality score; and a location determination module, configured to determine the location of the target merchant based on the location data set with the highest quality as represented by the comprehensive quality score.

[0006] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the merchant location data processing method of the first aspect.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the merchant location data processing method of the first aspect.

[0008] This application provides a method, apparatus, device, and storage medium for processing merchant location data. It can cluster location data sets from multiple channels for a target merchant, resulting in clusters for each location data set. Based on individual analysis of each location data set and data quality rules, a self-quality score characterizing the data quality of each location data set can be obtained. Based on the clusters of different location data sets, a cross-reference quality score characterizing the data quality reflected in the mutual reference between location data sets is obtained. Considering both the self-quality score and the cross-reference quality score, a comprehensive quality score characterizing the overall quality of the location data sets is obtained. By using location data from location data sets from channels with relatively higher overall quality to determine the merchant's location, the adverse effects of inconsistent and conflicting data quality from different channels on the accuracy of merchant positioning can be reduced, thereby improving the accuracy of merchant positioning. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating a merchant location data processing method provided in an embodiment of this application;

[0011] Figure 2 A schematic diagram illustrating an example of a clustering cluster provided in an embodiment of this application;

[0012] Figure 3 A flowchart illustrating a merchant location data processing method provided in another embodiment of this application;

[0013] Figure 4 A schematic diagram illustrating an example of a first convex polygon in a target cluster provided in an embodiment of this application;

[0014] Figure 5A schematic diagram illustrating an example of a location point network provided in an embodiment of this application;

[0015] Figure 6 This is a schematic diagram of the structure of a merchant location data processing device provided in an embodiment of this application;

[0016] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples. It should be noted that the acquisition, storage, use, and processing of information and data in the embodiments of this application are all authorized by users or relevant organizations and comply with the relevant provisions of national laws and regulations.

[0018] With the continuous development of information technology, more and more merchants' information services need to be based on their location. For merchants, location data can be obtained from multiple sources, and the accurate location of the merchant can be determined based on the location data from multiple sources. However, the quality of location data from different sources varies, and there may be conflicts between them, making it difficult to obtain the accurate location of the merchant even when location data from multiple sources is available, thus reducing the accuracy of merchant location positioning.

[0019] This application provides a method, apparatus, device, and storage medium for processing merchant location data. It first processes location data from multiple sources individually to obtain the quality of each source's location data. Then, it constructs a network based on the location data from multiple sources and, based on the network relationships, obtains the quality of cross-referencing between the location data from different sources. By comprehensively considering both the individual quality of each source and the quality of cross-referencing, it obtains the overall quality of the source location data. The merchant's location is determined based on the source with the best overall quality. This approach reduces the adverse effects of inconsistent and conflicting data quality from different sources on the accuracy of merchant positioning, improves the accuracy of merchant positioning, and preserves the precision of the location data, avoiding precision loss.

[0020] The first aspect of this application provides a merchant location data processing method, which can be applied to scenarios where merchants have two or more sources of location data. This merchant location data processing method can be executed by a merchant location data processing device, equipment, etc. Figure 1 A flowchart of a merchant location data processing method provided in an embodiment of this application is shown below. Figure 1 As shown, the merchant location data processing method may include steps S101 to S105.

[0021] In step S101, the location data set of each channel source of the target merchant is clustered to obtain a cluster of each location data set.

[0022] A target merchant will correspond to location data from two or more channel sources. A primary key can be configured for each location data point; the primary key identifies the merchant and allows filtering of location data corresponding to the target merchant. A location dataset includes location data; the number of location data points in a dataset is not limited, and a dataset can include one or more location data points. Location data can represent geographic locations; in some examples, location data may include latitude and longitude, for example, Global Positioning System (GPS) data or other data that can precisely pinpoint locations. Location data from each channel source of the target merchant can form a location dataset, and each location dataset can correspond one-to-one with a channel source. For example, for a target merchant, the location datasets from multiple channel sources can be implemented as a set J. <J1,J2,…,J n > where J1 represents the location data set of the target merchant's first channel source, J2 represents the location data set of the target merchant's second channel source, J n This represents the location data set from the nth channel source of the target merchant; location data set J n It can include m location data points, and these m location data points can be used by P. n1 P n2 ... P nm This means that each location data point can include latitude and longitude, such as location data P. nm It can be represented as <lat m , long m >,lat m Represents location data P nm latitude, long m Represents location data P nm Longitude.

[0023] In some examples, the source of location data may include two or more of the following: user terminals, merchant devices, registry authorities, and third-party organizations. The location data set from user terminals may include location data uploaded by the user terminals when the target merchant is used. The location data set from merchant devices may include location data uploaded by the merchant devices when the target merchant is used. The location data set from registry authorities may include location data provided by the target merchant during registration with the registry authority. The location data set from third-party organizations may include location data collected by third-party organizations that indicates the location of the target merchant.

[0024] Location data from different sources may be different. To facilitate the analysis of the characteristics of location data in location datasets, the location data in each location dataset can be clustered. Each location dataset can be clustered into one or more clusters, and each cluster includes at least one location data.

[0025] In some examples, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm can be used for clustering. As an unsupervised algorithm, DBSCAN can discover spatial clusters of arbitrary shapes and effectively handle noise points, and its clustering speed is relatively fast, making it suitable for clustering location data in location datasets in this application embodiment. A preset minimum clustering distance threshold and a minimum number of location data points within a cluster can be obtained. Based on the minimum clustering distance threshold and the minimum number of location data points within a cluster, the location data sets from each channel source of the target merchant are clustered to obtain a cluster for each location data set. The minimum clustering distance threshold is the minimum distance threshold required to divide different location data into two classes. If the location distance between two pieces of location data is greater than this minimum clustering distance threshold, the two pieces of location data can be divided into different clusters. The minimum number of location data points within a cluster is the minimum number of location data points within the cluster. The number of location data points must be greater than or equal to the minimum number of location data points within the cluster to form a cluster. For example, Figure 2 A schematic diagram illustrating an example of a clustering cluster provided in an embodiment of this application, as shown below. Figure 2 As shown, the location data in the location data set corresponding to a certain source channel, after clustering, yields three clusters: cluster A1, cluster A2, and cluster A3. In the above embodiment, when the location data includes latitude and longitude, the location distance between two location data can be the geographical distance, which can be calculated using the following formula (1):

[0026] d=R*arccos(sin(lat1)sin(lat2)cos(long1-long2)+cos(lat1)cos(lat2)) (1)

[0027] Where d is the distance between the two location data; R is the Earth's radius; lat1 is the latitude of the first location data; lat2 is the latitude of the second location data; long1 is the longitude of the first location data; and long2 is the longitude of the second location data.

[0028] In step S102, based on the cluster and the preset data quality rules, the quality score of each location data set is obtained.

[0029] Data quality rules can include the correspondence between the attribute parameters of clusters and their own quality scores. The attribute parameters of clusters obtained by clustering location data in a location dataset can be used to determine the own quality score of that location dataset within the data quality rules. The own quality score of a location dataset is determined through individual analysis of that location dataset and can characterize the data quality of the location dataset itself, i.e., the quality of the location data from different channels. Location data collected from various channels may be offset. The better the quality of a location dataset from a particular channel for a target merchant, the more clustered the location data should be. The own quality score of a location dataset can be determined by identifying the cluster containing the most location data within the corresponding cluster. The own quality score can be positively or negatively correlated with the data quality of the location dataset itself. For ease of explanation, this embodiment uses a positive correlation between the own quality score and the data quality of the location dataset itself; that is, a higher own quality score indicates higher data quality of the location dataset itself.

[0030] In step S103, based on the clustering, the cross-reference quality score of each location data set relative to other location data sets is obtained.

[0031] Different location datasets can be cross-referenced. Based on the relationships between clusters in different location datasets, a cross-reference quality score can be obtained relative to other location datasets. This cross-reference quality score characterizes the data quality reflected in cross-references between location datasets from different sources. When there is little conflict between different location datasets, the distances between clusters in different location datasets are very close; similarly, when there is significant conflict between different location datasets, the distances between clusters in different location datasets are larger. The distances between clusters in other location datasets and the clusters in this location dataset can also reflect the data quality of this location dataset to some extent. The cross-reference quality score of a location dataset relative to other location datasets can be determined based on these distances. The cross-reference quality score can be positively correlated with the data quality reflected by cross-reference between location datasets, or it can be negatively correlated with the data quality reflected by cross-reference between location datasets. For ease of explanation, this application embodiment uses the example of a positive correlation between the cross-reference quality score and the data quality reflected by cross-reference between location datasets, that is, the higher the cross-reference quality score, the higher the data quality reflected by cross-reference between location datasets.

[0032] The execution order of steps S102 and S103 is not limited here. Step S102 can be executed before step S103, or after step S103, or steps S102 and S103 can be executed simultaneously.

[0033] In step S104, the overall quality score of the location dataset is obtained based on its own quality score and the cross-reference quality score.

[0034] The overall quality score of a location dataset is obtained by combining its own quality score and the cross-reference quality scores of other location datasets. This overall quality score characterizes the overall quality of the location dataset. The overall quality score can be positively or negatively correlated with the overall quality of the location dataset. For ease of explanation, this embodiment illustrates a positive correlation between the overall quality score and the overall quality of the location dataset; that is, a higher overall quality score indicates a higher overall quality of the location dataset.

[0035] In some examples, a weighted algorithm can be used to calculate the overall quality score. The overall quality score of the location dataset can be obtained by weighting its own quality score, cross-reference quality scores, a first weight, and a second weight. The first weight corresponds to the own quality score, and the second weight corresponds to the cross-reference quality score. For example, the overall quality score of the location dataset can be calculated using the following formula (2):

[0036] S f =αS ori +βS avg (2)

[0037] Among them, S f The overall quality score; α is the first weight; S ori β is its own quality score; β is the second weight; S avg For mutual reference quality scores.

[0038] The values ​​of the first and second weights can be determined based on the data processing strategy. If the data quality of the location dataset itself is more important than the data quality reflected in the cross-references between location datasets, then the first weight is greater than the second weight. Conversely, if the data quality reflected in the cross-references between location datasets is more important than the data quality reflected in the cross-references between location datasets, then the first weight is less than the second weight. In some cases, only the data quality of the location dataset itself can be considered, in which case the first weight is 1 and the second weight is 0. In other cases, only the data quality reflected in the cross-references between location datasets can be considered, in which case the first weight is 0 and the second weight is 1.

[0039] In step S105, the location of the target merchant is determined based on the set of location data with the highest quality as represented by the comprehensive quality score.

[0040] When the overall quality score is positively correlated with the overall quality of the location dataset, the location dataset with the highest overall quality score can be selected. This dataset has the highest overall quality, and the center location can be calculated using the location data in it. This center location is then determined as the location of the target merchant. In some examples, the average value of the location data in the location dataset with the highest overall quality can be used to determine the location of the target merchant. In other examples, the center location of the cluster containing the most location data in the location dataset with the highest overall quality can be used to determine the location of the target merchant. In still other examples, the center location of all location data in the location dataset with the highest overall quality can be used to determine the location of the target merchant. Other methods that can determine the location of the target merchant based on the location dataset with the highest quality, as represented by the overall quality score, are also within the scope of protection of this application's embodiments and will not be elaborated upon here.

[0041] In this embodiment, location data sets from multiple channels for a target merchant can be clustered to obtain clusters for each location data set. Based on individual analysis of each location data set and data quality rules, a self-quality score characterizing the data quality of each location data set can be obtained. Based on the clusters of different location data sets, a cross-reference quality score characterizing the data quality reflected in the mutual reference between location data sets is obtained. Considering both the self-quality score and the cross-reference quality score, a comprehensive quality score characterizing the overall quality of the location data set is obtained. Using location data from a channel source with relatively better overall quality to determine the merchant's location reduces the adverse effects of inconsistent and conflicting data quality from different channels on the accuracy of merchant positioning, thereby improving the accuracy of merchant positioning. Furthermore, since the merchant's location is determined based on location data without using geohashing or other methods, the accuracy of the location data is preserved, avoiding loss of accuracy in merchant positioning.

[0042] In some embodiments, the location data set’s own quality score and cross-reference quality score can be determined based on the largest cluster in the clusters of the location data set. Figure 3 A flowchart illustrating a merchant location data processing method provided in another embodiment of this application. Figure 3 and Figure 1 The difference is that, Figure 1 Step S102 can be further refined as follows: Figure 3 Steps S1021 and S1022 in the process, Figure 1 Step S103 can be further refined as follows: Figure 3 Steps S1031 to S1033 in the process.

[0043] In step S1021, the target cluster containing the most location data is selected from the clusters of each location data set.

[0044] After clustering, a location dataset can be divided into at least one cluster. The number of location data points in each cluster can be counted, and the cluster containing the most location data points is determined as the target cluster. For example, if the location dataset has clusters like... Figure 2 As shown, Figure 2 The target cluster for the mid-location dataset is cluster A3.

[0045] In step S1022, the quality score of the location data set is determined based on the number of location data in the target cluster, the proportion of location data in the target cluster, and the data quality rules.

[0046] The proportion of location data in the target cluster is the ratio of the number of location data in the target cluster to the total number of location data in the location dataset, which can be represented by the Max Cluster Point Rate (MCPR). The number of location data in the target cluster can be represented by the Max Cluster Point Number (MCPN), and the number of location data in the location dataset can be represented by the All Cluster Point Number (ACPN). For example, the proportion of location data in the target cluster can be calculated according to the following formula (3):

[0047]

[0048] Data quality rules include the correspondence between the quantity of location data in the target cluster, the proportion of location data in the target cluster, and its own quality score. The more location data there is in the target cluster, and the larger the proportion of location data, the more concentrated the location data is in the location dataset, and correspondingly, the higher the data quality of the location dataset itself. In some examples, data quality rules may include the correspondence between the quantity of location data in the target cluster, the proportion of location data in the target cluster, its own quality level, and its own quality score. The quantity of location data in the target cluster can be divided into multiple quantity ranges, and the proportion of location data in the target cluster can be divided into multiple proportion ranges. Each quantity range and each proportion range corresponds to a quality score; the higher the lower limit of the quantity range and the lower limit of the proportion range, the higher the corresponding quality score. The correspondence between the quantity range, proportion range, and quality score can be set according to the scenario, requirements, experience, etc. For example, if the number of location data in the target cluster is greater than the first quantity threshold N1, and the proportion of location data in the target cluster is greater than the first proportion threshold R1, then its quality level is level one, and its quality score is S1. If the number of location data in the target cluster is greater than the second quantity threshold N2 and less than or equal to the first quantity threshold N1, and the proportion of location data in the target cluster is greater than the second proportion threshold R2 and less than or equal to the first proportion threshold R1, then its quality level is level two, and its quality score is S2. The second quantity threshold N2 is less than the first quantity threshold N1, the second proportion threshold R2 is less than the first proportion threshold R1, and its quality score S2 is less than its quality score S1. If the number of location data in the target cluster is greater than the third quantity threshold N1, then its quality level is level one, and its quality score is S2. The second quantity threshold N2 is less than the first quantity threshold N1, the second proportion threshold R2 is less than the first proportion threshold R1, and its quality score S2 is less than its quality score S1. If the value N3 is less than or equal to the second quantity threshold N2, and the proportion of location data in the target cluster is greater than the third proportion threshold R3 and less than or equal to the second proportion threshold R2, then the self-quality level is the third level, and the self-quality score is S3. The third quantity threshold N3 is less than the second quantity threshold N2, the third proportion threshold R3 is less than the second proportion threshold R2, and the self-quality score S3 is less than the self-quality score S2. In other cases, the self-quality level is the fourth level, and the self-quality score is S4. The self-quality score S4 is less than the self-quality score S3. For example, if the number of location data in the target cluster is less than the first quantity threshold N1, but the proportion of location data in the target cluster is greater than the first proportion threshold R1, the corresponding self-quality score is S4.

[0049] The above data quality rules can be expressed as follows (4):

[0050]

[0051] Among them, S oriThe first quantity threshold N1, the second quantity threshold N2, the third quantity threshold N3, the first proportion threshold R1, the second proportion threshold R2, the third proportion threshold R3, the self-quality score S1, the self-quality score S2, the self-quality score S3, and the self-quality score S4 can be set according to the scenario, requirements, experience, etc. For example, the first quantity threshold N1 is 20, the first proportion threshold R1 is 70%, and the self-quality score S1 is 10.

[0052] In step S1031, the percentage of location data in the target cluster that contains the most location data in the location data set is obtained.

[0053] For details regarding the proportion of location data in the target cluster, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0054] In step S1032, the center location data of the location data set is determined based on the proportion of location data in the target cluster and a preset proportion threshold.

[0055] The percentage threshold can be used to determine whether a target merchant's location is within a target cluster. The value of the percentage threshold can be set based on the scenario, requirements, experience, etc. In some examples, the percentage threshold here can be equal to the lower limit of the percentage range corresponding to the highest self-quality score in the data quality rules of the above embodiments; for example, the percentage threshold can be 70%. If the percentage of location data in the target cluster is greater than or equal to the percentage threshold, the target merchant's location can be considered to be within the target cluster; if the percentage of location data in the target cluster is less than the percentage threshold, the probability that the target merchant's location is within the target cluster is considered low. Based on the probability that the target merchant's location is within the target cluster, the central location data of the location data set can be determined from which location data in the location data set to obtain the central location data of the location data set. The central location data of the location data set is used to characterize the center of the location data in the location data set.

[0056] In some examples, when the proportion of location data in the target cluster is greater than or equal to the proportion threshold, a first convex polygon is formed based on the location data located on the periphery of the target cluster, and the centroid of the first convex polygon is calculated as the center location data. If the proportion of location data in the target cluster is greater than or equal to the proportion threshold, the location of the target merchant is likely to be located in the target cluster. When determining the center location data, other location data in the location data set besides the target cluster can be ignored. The location data on the periphery of the target cluster are connected end to end to form a first convex polygon, and the centroid of the first convex polygon is used as the center location data. The centroid of the first convex polygon can be calculated based on the area of ​​the first convex polygon and the location data of the vertices of the first convex polygon. For example, the centroid of the first convex polygon, i.e., the center location data, can be calculated according to the following formulas (5) and (6):

[0057]

[0058]

[0059] Here, let the first convex polygon have N+1 vertices, lat i Let lat be the latitude of the i-th vertex of the first convex polygon. i+1 Let be the latitude of the (i+1)th vertex of the first convex polygon, long i Let be the longitude of the i-th vertex of the first convex polygon, long i+1 Let x be the longitude of the (i+1)th vertex of the first convex polygon, A be the area of ​​the first convex polygon, and x be the longitude of the (i+1)th vertex. c Let y be the latitude of the centroid of the first convex polygon. c Let be the longitude of the centroid of the first convex polygon.

[0060] For example, Figure 4 A schematic diagram illustrating an example of a first convex polygon in a target cluster provided in this application embodiment, as shown below. Figure 4 As shown, the location data set includes cluster A4 and cluster A5. Each point in cluster A4 and cluster A5 represents a location data. Cluster A5 is the target cluster. The location data B1 to B8 of the outer periphery of the target cluster A5 can be connected to form a first convex polygon. The center location data corresponding to the first convex polygon can be calculated according to the above formulas (5) and (6).

[0061] In some examples, when the proportion of location data in the target cluster is less than the proportion threshold, a second convex polygon is formed based on the location data located on the periphery of the target cluster, or based on the location data located on the periphery of the location data set, and the centroid of the second convex polygon is calculated as the center location data. When the proportion of location data in the target cluster is less than the proportion threshold, it can be considered that the location of the target merchant is relatively unlikely to be in the target cluster. When determining the center location data, other location data in the location data set besides the target cluster can be ignored, and the location data on the periphery of the target cluster can be connected end to end to form a second convex polygon, and the centroid of the second convex polygon can be used as the center location data. Alternatively, when determining the center location data, the location data located on the periphery of all location data in the location data set can be connected end to end to form a second convex polygon, and the centroid of the second convex polygon can be used as the center location data. The centroid of the second convex polygon, i.e., the center location data, can also be calculated according to the above formulas (5) and (6). Correspondingly, when calculating the centroid of the second convex polygon, lat i Let lat be the latitude of the i-th vertex of the second convex polygon. i+1 Let be the latitude of the (i+1)th vertex of the second convex polygon, long i Let longitude be the longitude of the i-th vertex of the second convex polygon. i+1 Let x be the longitude of the (i+1)th vertex of the second convex polygon, A be the area of ​​the second convex polygon, and x be the longitude of the (i+1)th vertex. c Let y be the latitude of the centroid of the second convex polygon. c Let be the longitude of the centroid of the second convex polygon.

[0062] In step S1033, based on the location distance between the central location data of the target merchant's location data set, the cross-reference quality score of each location data set relative to other location data sets is obtained.

[0063] The center location data of each location dataset indicates the center of that dataset, and the distance between the center locations of different location datasets is the distance between the centers of the datasets. The closer the centers of different location datasets for the same target merchant are, the higher the data quality of the cross-reference between the different location datasets, and correspondingly, the higher the cross-reference quality score. Based on the location distance between the center locations of the target merchant's location datasets, the cross-reference quality score of each location dataset can be obtained compared to other location datasets. The data quality represented by the cross-reference quality score is negatively correlated with the location distance; the closer the location distance, the higher the data quality represented by the cross-reference quality score. The average cross-reference quality score is then calculated and used as the cross-reference quality score.

[0064] For ease of processing, the central location data of each location dataset of the target merchant can be taken as a point, and the central location data of each location dataset of the target merchant can be connected in pairs to form a LocationPoint Net (LPN). The length of the edge formed by connecting each pair of points in the LocationPoint Net is the location distance. The calculation of the location distance can be found in Equation (1) above, and will not be repeated here. For example, Figure 5 A schematic diagram illustrating an example of a location point network provided in an embodiment of this application, as shown below. Figure 5 As shown, this includes center position data C1, C2, and C3. For center position data C1, there are two connecting edges: the edge between center position data C1 and center position data C2, and the edge between center position data C1 and center position data C3. The length of the edge between center position data C1 and center position data C2 is the position distance between center position data C1 and center position data C2, and the length of the edge between center position data C1 and center position data C3 is the position distance between center position data C1 and center position data C3.

[0065] The calculation rules for the location distance and reference quality score between the center location data of two location datasets can be preset, and the location distance between the center location data of each pair of location datasets can be converted into a reference quality score.

[0066] In some examples, the location distance can be pre-divided into multiple distance ranges, with higher reference quality scores corresponding to locations within larger numerical ranges. For example, when the location distance is less than or equal to the location threshold d1, the reference quality score is P1; when the location distance is greater than the location threshold d1 but less than or equal to the location threshold d2, the reference quality score is P2; when the location distance is greater than the location threshold d2, the reference quality score is P3; when the location threshold d1 is less than the location threshold d2, the reference quality score P1 is greater than the reference quality score P2, and the reference quality score P2 is greater than the reference quality score P3; correspondingly, the above calculation rules can be obtained according to the following formula (7):

[0067]

[0068] Among them, S d d represents the reference quality score; d represents the location distance. The values ​​of location thresholds d1, d2, reference quality scores P1, P2, and P3 can be set according to the scenario, requirements, and experience. For example, location threshold d1 can be 50 meters, location threshold d2 can be 500 meters, reference quality score P1 can be 100 points, reference quality score P2 can be 80 points, and reference quality score P3 can be 60 points.

[0069] In other examples, the computation rule can be implemented using a function. For example, the computation rule can be implemented using the following equation (8):

[0070]

[0071] Among them, S d d represents the reference mass fraction; d represents the location distance.

[0072] The aforementioned reference quality score reflects the data quality of two location datasets in mutual reference. However, for a location dataset, it may be mutually referenced with more than one other location dataset. The mutual reference quality score corresponding to that location dataset can be obtained based on the reference quality scores of that location dataset with each of the other location datasets. Specifically, for each location dataset, the mutual reference quality score corresponding to that location dataset can be calculated according to the following formula (9):

[0073]

[0074] Here, it is assumed that the center location data of this location dataset is connected to the center location data of N other location datasets by edges; S avg S represents the cross-reference quality score of the dataset at this location; di This is the reference quality score between this location dataset and another location dataset of the i-th location.

[0075] The second aspect of this application provides a merchant location data processing device. Figure 6 This is a schematic diagram of the structure of a merchant location data processing device provided in an embodiment of this application, as shown below. Figure 6 As shown, the merchant location data processing device 200 may include a clustering module 201, a first score determination module 202, a second score determination module 203, a comprehensive score determination module 204, and a location determination module 205.

[0076] The clustering module 201 can be used to cluster the location data set of each channel source of the target merchant to obtain the cluster of each location data set, where the location data set includes location data.

[0077] In some examples, the source of the channel includes two or more of the following: user terminal, merchant equipment, registration management agency, and third-party agency.

[0078] The first score determination module 202 can be used to obtain the quality score of each location data set based on clustering and preset data quality rules.

[0079] The second score determination module 203 can be used to obtain the cross-reference quality score of each location data set relative to other location data sets based on clustering.

[0080] The comprehensive score determination module 204 can be used to obtain the comprehensive quality score of the location dataset based on its own quality score and the cross-reference quality scores.

[0081] The location determination module 205 can be used to determine the location of a target merchant based on the set of location data with the highest quality as represented by the comprehensive quality score.

[0082] In this embodiment, location data sets from multiple channels for a target merchant can be clustered to obtain clusters for each location data set. Based on individual analysis of each location data set and data quality rules, a self-quality score characterizing the data quality of each location data set can be obtained. Based on the clusters of different location data sets, a cross-reference quality score characterizing the data quality reflected in the mutual reference between location data sets is obtained. Considering both the self-quality score and the cross-reference quality score, a comprehensive quality score characterizing the overall quality of the location data set is obtained. Using location data from a channel source with relatively better overall quality to determine the merchant's location reduces the adverse effects of inconsistent and conflicting data quality from different channels on the accuracy of merchant positioning, thereby improving the accuracy of merchant positioning. Furthermore, since the merchant's location is determined based on location data without using geohashing or other methods, the accuracy of the location data is preserved, avoiding loss of accuracy in merchant positioning.

[0083] In some embodiments, the first score determination module 202 described above may be specifically used to: select the target cluster containing the most location data from the clusters of each location data set; determine the quality score of the location data set itself based on the number of location data in the target cluster, the proportion of location data in the target cluster, and data quality rules, wherein the data quality rules include the correspondence between the number of location data in the target cluster, the proportion of location data in the target cluster, and the quality score itself.

[0084] In some embodiments, the second score determination module 203 described above may be specifically used to: obtain the proportion of the number of location data in the target cluster that contains the most location data in the location data set; determine the center location data of the location data set based on the proportion of the number of location data in the target cluster and a preset proportion threshold; and obtain the mutual reference quality score of each location data set relative to other location data sets based on the location distance between the center location data of the target merchant's location data set.

[0085] In some examples, the second score determination module 203 described above can be specifically used to: when the proportion of the number of location data in the target cluster is greater than or equal to the proportion threshold, form a first convex polygon based on the location data located on the periphery of the target cluster, and calculate the centroid of the first convex polygon as the center location data; when the proportion of the number of location data in the target cluster is less than the proportion threshold, form a second convex polygon based on the location data located on the periphery of the target cluster, or based on the location data located on the periphery in the location data set, and calculate the centroid of the second convex polygon as the center location data.

[0086] In some examples, the second score determination module 203 described above can be specifically used to: obtain a reference quality score for each location data set of the target merchant and other location data sets based on the location distance between the central location data of the target merchant's location data set, wherein the data quality represented by the reference quality score is negatively correlated with the location distance; calculate the average of the reference quality scores of the location data set and other location data sets, and use the average as the mutual reference quality score.

[0087] In some embodiments, the above-mentioned comprehensive score determination module 204 can be used to: perform a weighted calculation based on the location data set's own quality score, cross-reference quality score, first weight, and second weight to obtain the comprehensive quality score of the location data set, where the first weight corresponds to the own quality score and the second weight corresponds to the cross-reference quality score.

[0088] In some embodiments, the clustering module 201 may be specifically used to: obtain a preset minimum clustering distance threshold and a minimum number of location data within a cluster; and cluster the location data set of each channel source of the target merchant according to the minimum clustering distance threshold and the minimum number of location data within a cluster to obtain a cluster of each location data set.

[0089] A third aspect of this application also provides an electronic device. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 300 includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302.

[0090] In some examples, the processor 302 described above may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that may be configured to implement the embodiments of this application.

[0091] Memory 301 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the merchant location data processing method according to embodiments of this application.

[0092] The processor 302 runs a computer program corresponding to the executable program code by reading the executable program code stored in the memory 301, so as to implement the merchant location data processing method in the above embodiments.

[0093] In some examples, the electronic device 300 may also include a communication interface 303 and a bus 304. For example, Figure 7 As shown, the memory 301, processor 302, and communication interface 303 are connected through bus 304 and complete communication with each other.

[0094] The communication interface 303 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application. Input devices and / or output devices can also be connected through the communication interface 303.

[0095] Bus 304 includes hardware, software, or both, that couples components of electronic device 300 together. For example, and not limitingly, bus 304 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 304 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0096] A fourth aspect of this application also provides a computer-readable storage medium storing computer program instructions. When executed by a processor, these computer program instructions can implement the merchant location data processing method described in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.

[0097] This application provides a computer program product in which the instructions are executed by the processor of an electronic device, causing the electronic device to perform the merchant location data processing method in the above embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0098] It should be clarified that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. For the device embodiments, equipment embodiments, and computer-readable storage medium embodiments, the relevant parts can be referred to the description section of the method embodiments. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application. Furthermore, for the sake of brevity, detailed descriptions of known methods and techniques are omitted here.

[0099] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0100] Those skilled in the art will understand that the above embodiments are exemplary and not restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, specification, and claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other means or steps; the quantifier "a" does not exclude a plurality; the terms "first" and "second" are used to identify names and not to indicate any particular order. No reference numerals in the claims should be construed as limiting the scope of protection. The functionality of multiple parts appearing in the claims can be implemented by a single hardware or software module. The appearance of certain technical features in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A method for processing merchant location data, characterized in that, include: Clustering is performed on the location data set of each channel source for the target merchant to obtain a cluster of each location data set, wherein the location data set includes location data and the location data set corresponds one-to-one with the channel source; From the clusters of each location data set, select the target cluster that includes the largest number of location data. The quality score of the location data set is determined based on the number of location data in the target cluster and the proportion of the number of location data in the target cluster. The proportion of the number of location data in the target cluster is the ratio of the number of location data in the target cluster to the number of location data in the location data set. Based on the proportion of the location data in the target cluster and a preset proportion threshold, the center location data of the location data set is determined; Based on the location distance between the central location data of the location data set of the target merchant, a cross-reference quality score is obtained for each location data set relative to other location data sets, and the cross-reference quality score is negatively correlated with the location distance between the central location data of the location data set of the target merchant; The overall quality score of the location dataset is obtained based on its own quality score and the cross-reference quality score. The location of the target merchant is determined based on the set of location data with the highest quality as represented by the comprehensive quality score.

2. The method according to claim 1, characterized in that, The step of determining the quality score of the location data set based on the quantity of location data in the target cluster and the proportion of location data in the target cluster includes: Based on the quantity of location data in the target cluster, the proportion of location data in the target cluster, and data quality rules, the self-quality score of the location data set is determined. The data quality rules include the correspondence between the quantity of location data in the target cluster, the proportion of location data in the target cluster, and the self-quality score.

3. The method according to claim 1, characterized in that, The step of determining the center location data of the location data set based on the proportion of the location data in the target cluster and a preset proportion threshold includes: If the proportion of the location data in the target cluster is greater than or equal to the proportion threshold, a first convex polygon is formed based on the location data located on the periphery of the target cluster, and the centroid of the first convex polygon is calculated as the center location data. If the proportion of the location data in the target cluster is less than the proportion threshold, a second convex polygon is formed based on the location data located on the periphery of the target cluster, or based on the location data located on the periphery of the location data set, and the centroid of the second convex polygon is calculated as the center location data.

4. The method according to claim 1, characterized in that, The method of obtaining the cross-reference quality score of each location data set relative to other location data sets based on the location distance between the central location data of the location data set of the target merchant includes: Based on the location distance between the central location data of the location data set of the target merchant, a reference quality score is obtained for each location data set of the target merchant and other location data sets. The data quality represented by the reference quality score is negatively correlated with the location distance. Calculate the average of the reference quality scores of the location dataset and other location datasets, and use the average value as the mutual reference quality score.

5. The method according to claim 1, characterized in that, Based on the location dataset's own quality score and the cross-referenced quality score, a comprehensive quality score for the location dataset is obtained, including: The overall quality score of the location data set is obtained by weighting the self-quality score, the cross-reference quality score, the first weight, and the second weight. The first weight corresponds to the self-quality score, and the second weight corresponds to the cross-reference quality score.

6. The method according to claim 1, characterized in that, The clustering of location data sets for each channel source of the target merchant to obtain clusters for each location data set includes: Obtain the preset minimum clustering distance threshold and the minimum number of location data within each cluster; Based on the minimum clustering distance threshold and the minimum number of location data within the cluster, the location data sets from each channel source of the target merchant are clustered to obtain a cluster for each location data set.

7. The method according to any one of claims 1 to 6, characterized in that, The sources of these channels include two or more of the following: User terminals, merchant equipment, registration management agencies, and third-party organizations.

8. A merchant location data processing device, characterized in that, include: The clustering module is used to cluster the location data sets of each channel source of the target merchant to obtain the cluster of each location data set, wherein the location data set includes location data and the location data set corresponds one-to-one with the channel source; The first score determination module is used to select the target cluster that contains the most location data from the clusters of each location data set; and to determine the quality score of the location data set itself based on the number of location data in the target cluster and the proportion of the number of location data in the target cluster, wherein the proportion of the number of location data in the target cluster is the ratio of the number of location data in the target cluster to the number of location data in the location data set. The second score determination module is used to determine the center position data of the location data set based on the proportion of the location data in the target cluster and a preset proportion threshold. Based on the location distance between the central location data of the location data set of the target merchant, a cross-reference quality score is obtained for each location data set relative to other location data sets, and the cross-reference quality score is negatively correlated with the location distance between the central location data of the location data set of the target merchant; The comprehensive score determination module is used to obtain the comprehensive quality score of the location data set based on the location data set's own quality score and the cross-reference quality score; The location determination module is used to determine the location of the target merchant based on the location data set with the highest quality as represented by the comprehensive quality score.

9. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the merchant location data processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the merchant location data processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Travel recommendation method, computer device and readable storage medium

    CN110598778A

  • Method and device for identifying merchant position and electronic equipment

    CN110969483A