Factory Address Verification Method, Device and Electronic Equipment

By collecting and converging plant address information using the first data source and multiple second data sources, determining the equivalent and confidence of the address, the problems of inconsistency in plant address and insufficient data source coverage are solved, and more accurate and comprehensive plant address verification is achieved.

CN114168647BActive Publication Date: 2025-06-27ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111329402.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-06-27
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

In the B2B e-commerce system, the registered address of the factory is inconsistent with the actual business address, which leads to the inability to complete the factory visit through the registered address alone, and requires the collection of more real and valid factory address information. However, the existing data sources do not cover and update the factory in time, resulting in inaccurate address information.

Method used

By determining the factory to be verified, the registered address of the factory is obtained using the first data source, and address information is collected from multiple second data sources, address equivalent to the first data source address is determined for some or all of the second data sources, address amplification is performed, and through multi-data source fusion processing, the confidence that the addresses provided by each data source belong to the real address is determined.

Benefits of technology

Improve the accuracy and coverage of factory address information, ensure that more factories can obtain more accurate address information, and support more efficient factory visits and order acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168647B_ABST
    Figure CN114168647B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a factory address verification method, apparatus, and electronic device. Among them, the method includes: determining a plurality of factories to be verified for addresses, determining first addresses respectively corresponding to the plurality of factories through a first data source, and collecting second addresses for the plurality of factories from a plurality of second data sources; for some or all of the plurality of second data sources, determining, from the first addresses respectively corresponding to the plurality of factories, first addresses equivalent to the second addresses provided by the second data sources; using the equivalent first addresses to amplify the second addresses provided by the second data sources; based on the address amplification results corresponding to each second data source, performing multi-data source fusion processing to determine the confidence levels that the addresses provided or amplified by the multiple data sources for each factory belong to real addresses. Through the embodiments of the present application, it is beneficial to obtain more accurate address information for more factories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of information processing, and particularly to a method, apparatus, and electronic device for verifying a factory address. Background Art

[0002] In an e-commerce system with a B2B (Business-to-Business) model, its seller users include factories. Specifically, a factory can register as a premium member in the system to preferentially view information such as purchase inquiry information and contact information posted by buyers in the system, enabling the seller to have the opportunity to obtain more orders preferentially.

[0003] In order to enable more factories to register as premium members, the staff in the system can use the method of on-site visits to factories (usually referred to as the "ground promotion" method) to do so. This requires establishing a factory address information database to provide the staff with the address information of factories, so that the staff can complete the visit work based on the address information of specific factories.

[0004] Generally, given the factory name, the registered address information of the factory can be obtained by querying business administration information, etc. However, in reality, the registered address often does not match the actual business address. Therefore, it is often impossible to complete the visit to the factory only through the registered address, and more real and effective factory address information needs to be collected.

[0005] In practical applications, the address information of factories can be obtained through various other data sources. However, the coverage of factories by data sources other than the registered address information is often insufficient. For example, some digital map information systems may include the POI (Point of Interest) address information of some factories, enabling the address information of factories to be obtained by entering the factory name in the system or calling the API (Application Programming Interface) provided by the system with the factory name as a parameter. Since the specific POI service usually needs to be paid for, the accuracy of the POI addresses provided by this data source is often relatively high. However, also because the map information system will provide specific POI address information only when the factory pays, there is a problem of low coverage of factories. For example, assuming there are a total of ten million factories, the map information system may be able to provide POI addresses for only hundreds of thousands of them, and so on.

[0006] In addition to the POI addresses in the map information system, there can also be some other data sources. For example, some industrial parks may provide the names and address information of the factories within the park, and so on. However, since the information provided by these data sources does not require information updates compulsorily, the address information provided by these data sources may not only have the aforementioned drawback of low coverage, but also may have the situation of untimely updates. That is, if a factory relocates or the like, and the information provided by the data source is not updated in time, the address information provided by the data source will also be inaccurate, and so on.

[0007] Therefore, how to obtain true and effective address information about factories has become a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0008] The present application provides a factory address verification method, device and electronic device, which is beneficial to obtaining more accurate address information for more factories.

[0009] The present application provides the following solutions:

[0010] A factory address verification method includes:

[0011] Determine a plurality of factories to be verified for addresses, determine the first addresses respectively corresponding to the plurality of factories through a first data source, and collect second addresses for the plurality of factories from a plurality of second data sources;

[0012] For some or all of the plurality of second data sources, determine, from the first addresses respectively corresponding to the plurality of factories, the first addresses equivalent to the second addresses provided by the second data sources;

[0013] Utilize the equivalent first addresses to amplify the second addresses provided by the second data sources;

[0014] Based on the address amplification results corresponding to each second data source, perform multi-data source fusion processing to determine the confidence levels that the addresses provided or amplified by the multiple data sources for each factory belong to true addresses.

[0015] Among them, the determination of the first addresses equivalent to the addresses provided by the second data sources includes:

[0016] For one of the second data sources, from the plurality of factories, determine the first addresses of some factories that can provide the second addresses by the second data source as the sample data of the second data source, and determine the distance between the first address and the second address of the same factory. If the distance is less than a first target threshold, determine the first address of the corresponding factory as a positive sample, otherwise as a negative sample;

[0017] Determine each first address within the area where the first addresses are densely distributed as a cluster according to the distances between pairs of the first addresses corresponding to the multiple factories;

[0018] By analyzing the first addresses within the multiple obtained clusters respectively, filter out the target cluster, and determine each first address in the target cluster as the first address equivalent to the second address provided by the second data source.

[0019] Among them, the step of determining each first address within the area where the first addresses are densely distributed as a cluster according to the distances between pairs of the first addresses corresponding to the multiple factories includes:

[0020] Within the same target zoning range, calculate the distances between pairs of the first addresses corresponding to the multiple factories, and determine each first address within the area where the first addresses are densely distributed within the same target zoning range as a cluster.

[0021] Among them, the step of calculating the distances between pairs of the first addresses corresponding to the multiple factories within the same target zoning range includes:

[0022] Map the boundary of the target zoning range on the sphere into the boundary on the plane, and approximately calculate the distances between pairs of the first addresses corresponding to the multiple factories through local Euclidean coordinates.

[0023] Among them, the step of analyzing the first addresses within the multiple obtained clusters respectively includes:

[0024] For one of the clusters, predict the number of the equivalent first addresses included in the cluster according to the number of the first addresses included in the cluster, the number of the first addresses belonging to the sample data among them, and the number of the positive samples among them. If the proportion of the number of the equivalent first addresses included in the cluster exceeds the second target threshold, determine it as the target cluster.

[0025] Among them, the step of predicting the number of the equivalent first addresses included in the cluster includes:

[0026] Based on the confidence estimation algorithm of the hypergeometric distribution, generate a function with the number of the first addresses included in the cluster, the number of the first addresses belonging to the sample data among them, and the number of the positive samples among them as known parameters and the number of the equivalent first addresses included in the cluster as an unknown parameter, and predict the number of the equivalent first addresses based on the method of the confidence interval.

[0027] Among them, the second data source includes the point of interest (POI) address data source provided by the map information system, and the second address includes the POI address;

[0028] The first address equivalent to the second address provided by the second data source includes:

[0029] The first address with the same confidence level as the POI address;

[0030] If the proportion of the number of the equivalent first addresses included in the cluster exceeds the second target threshold, determine the area corresponding to the cluster as an industrial park, determine the cluster as the target cluster, and determine each first address in the target cluster as the first address equivalent to the POI address provided by the POI address data source.

[0031] Among them, the multi-data-source fusion processing based on the address amplification results corresponding to each second data source includes:

[0032] Construct an algorithm model for multi-data-source fusion to predict the real address of a factory according to the addresses and / or amplified addresses provided by multiple data sources for the same factory, and determine the confidence level of the address provided by the data source for the factory according to the distance between the address provided by the data source for a certain factory and the prediction result of the real address corresponding to the factory.

[0033] Among them, the algorithm model includes probability parameters for characterizing the reliability of the addresses provided by each data source;

[0034] The prediction of the real address of the factory includes:

[0035] Perform multiple rounds of iteration on the algorithm model, and update the probability parameters during each iteration;

[0036] After the algorithm converges, determine the confidence level that the address provided or amplified by each data source for each factory belongs to the real address according to the output result of the corresponding algorithm model.

[0037] A factory address verification device includes:

[0038] An address collection unit for determining multiple factories to be verified for addresses, determining the first addresses corresponding to the multiple factories through a first data source, and collecting second addresses for the multiple factories from multiple second data sources;

[0039] An equivalent address determination unit for determining, for some or all of the multiple second data sources, the first addresses equivalent to the second addresses provided by the second data sources from the first addresses corresponding to the multiple factories respectively;

[0040] An address amplification unit for amplifying the second addresses provided by the second data sources by using the equivalent first addresses;

[0041] A multi-data source fusion unit is used to perform multi-data source fusion processing based on the address amplification results corresponding to each second data source, so as to determine the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses.

[0042] A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, the steps of the method described in any one of the foregoing are implemented.

[0043] An electronic device includes:

[0044] One or more processors; and

[0045] A memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the foregoing are executed.

[0046] According to the specific embodiments provided by the present application, the present application discloses the following technical effects:

[0047] Through the embodiments of the present application, for factories that need to verify addresses, address information can be collected for specific factories from multiple data sources, including determining the first addresses corresponding to multiple factories through the first data source, and collecting second addresses for the multiple factories from multiple second data sources. Then, for some or all of the multiple second data sources, the first addresses equivalent to the second addresses provided by the second data sources can be determined from the first addresses corresponding to the multiple factories respectively, and the second addresses provided by the second data sources can be amplified by using the equivalent first addresses. Then, based on the address amplification results corresponding to each second data source, multi-data source fusion processing can be performed to determine the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses. In this way, the potential connections between different data sources can be explored, and the first addresses provided by the first data source with a relatively high coverage can be used to amplify the second addresses provided by the second data source, thereby improving the coverage of the second data source for factories. When performing multi-data source fusion processing on this basis, it is beneficial to obtain more accurate address information for more factories.

[0048] Of course, it is not necessary for any product implementing the present application to achieve all the above advantages simultaneously. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0050] Figure 1 is a schematic diagram of the system architecture provided by the embodiments of the present application;

[0051] Figure 2 is a flowchart of the method provided by the embodiments of the present application;

[0052] Figure 3 is a schematic diagram based on the confidence interval provided by the embodiments of the present application;

[0053] Figure 4 is a schematic diagram of the device provided by the embodiments of the present application;

[0054] Figure 5 is a schematic diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.

[0056] In the embodiments of the present application, since the address information of the factory can be obtained through multiple data sources, it is possible to consider verifying the authenticity of the specific address by fusing the addresses provided by multiple data sources, and then selecting one address that can be used as the real address from the multiple addresses provided by multiple data sources for the same factory.

[0057] However, the prerequisite for data fusion is that for the same factory, there are multiple data sources providing address information for it. Otherwise, if there is only one data source providing address information for a factory, for example, only the registered address of a certain factory can be queried and other data sources cannot provide address information for this factory, then multi-data-source fusion will not even be possible. In addition to the data source providing the registered address, other data sources all have more or less low coverage in terms of the coverage of the factory address. This leads to the situation that if the address collection results of each data source are directly used for multi-data-source fusion, the coverage of the prediction result for the factory will also be relatively low, and it is difficult to achieve the preset goals in terms of the number of factories with verified addresses and the accuracy rate.

[0058] In view of the above situation, in the embodiments of the present application, first, the potential connections between the addresses provided by different data sources can be mined to find addresses equivalent to a specific data source (for example, having the same or comparable validity degree as the address provided by a certain data source, etc.), and these addresses are used to amplify the addresses provided by the data sources, that is, to increase the coverage of a single data source for the factory address. Then, based on the amplified results of each data source, multi-data source fusion is performed to achieve the coverage of the specific prediction result for the factory.

[0059] Among them, regarding the potential connections between the addresses provided by different data sources, it mainly refers to the potential connections between the registered address of the factory and the addresses provided by other data sources. Since the coverage of the registered address of a specific factory is often the highest, even up to 100%, therefore, this registered address can be used to amplify the addresses provided by other data sources (for the convenience of distinction, in the embodiments of the present application, the data source providing the registered address is referred to as the first data source, and other data sources are the second data sources). That is to say, in this way, assuming that a certain data source fails to provide an address for a certain factory, but by mining the potential relationship between the registered address and the address provided by this data source, it can be found that the registered address of this factory is equivalent to the address provided by this data source, then the registered address of this factory can be regarded as the address provided by this data source to achieve the amplification of the address provided by this data source. For example, for the POI address data source provided by the map information system, the POI addresses provided by it are usually considered valid addresses. However, the POI addresses can only cover some factories. If it is found through analysis that although the POI address data source cannot provide a POI address for a certain factory, but the probability that the registered address of this factory belongs to a valid address is also relatively high, then at this time, the registered address of this factory can also be regarded as a POI address and added to the addresses provided by the POI address data source to increase the number of valid addresses, and then multi-data source fusion is performed, etc.

[0060] In specific implementation, for the second data source that specifically needs to be amplified, in order to determine which registered addresses can become equivalent addresses of the addresses provided by this second data source, there are various ways. For example, in the embodiments of the present application, first, the factories for which this second data source can provide addresses can be determined, and the registered addresses of these factories are used as the sample data of this second data source. Then, according to the distance between the registered address of a certain factory and the address provided by this second data source for this factory, some positive samples and some negative samples can be determined.

[0061] In addition, pairwise distance calculations can be performed for the registered addresses of all factories, and then clustering can be carried out according to the density of the registered addresses of the factories. For example, each registered address within the area where the registered addresses are densely distributed can be determined as a cluster. That is to say, in practical applications, many factories may be characterized by clustered distribution within a certain industrial park or other areas. Therefore, if the registered addresses of the factories with clustered distribution can be determined first and grouped into clusters. After that, the clusters are analyzed to determine whether they are industrial parks. If so, the probability that all the registered addresses of the factories within the cluster belong to valid addresses will be relatively high. Furthermore, they can be used as equivalent addresses of data sources that can also provide relatively high valid addresses to achieve the amplification of the addresses provided by the data source.

[0062] Specifically, when analyzing the cluster, the number of factories included in the cluster, the number of registered addresses that belong to the sample data, and the number of positive samples can also be used to predict the number of equivalent registered addresses included in the cluster. If the proportion of the number of equivalent registered addresses included in the cluster exceeds a certain threshold, it is determined as the target cluster, and all the registered addresses in the target cluster are regarded as addresses equivalent to the addresses provided by the data source.

[0063] After the amplification of the addresses provided by the specific second data source is completed, multi-data source fusion processing can be carried out. Among them, an algorithm model for multi-data source fusion can be constructed (for example, a probability model based on distance-sensitive maximum likelihood estimation) to be used to predict the real address (here referring to the longitude and latitude of the real address) of the factory according to the addresses and / or amplified addresses provided by multiple data sources for the same factory, and determine the confidence level of the address provided by the data source for the factory according to the distance between the address provided by the specific data source for a factory and the prediction result of the real address corresponding to the factory. That is to say, the input of the algorithm model can be the amplified results corresponding to each data source, and the output can be the confidence levels of each address provided by each data source. Furthermore, the confidence levels corresponding to the addresses provided or amplified by multiple data sources for the same factory can be selected, and one that can represent the real address of the factory can be chosen, and so on.

[0064] Among them, in the process of specifically performing multi-data fusion, specific probability parameters can also be introduced into the algorithm model, and these probability parameters can be used to characterize the reliability of the addresses provided by the specific data source as a whole. In the process of fusion, multiple rounds of iteration can be carried out. In each round of iteration, the specific probability parameters can be updated so that they can more realistically characterize the reliability of the addresses provided by the corresponding data source. After the algorithm converges, the output result of the algorithm can be determined as the confidence levels of each address provided or amplified by each data source as the real address of the corresponding factory.

[0065] From the perspective of the system architecture, referring to Figure 1 , the embodiments of the present application can be used as a tool for verifying the addresses of factories. The tool can be divided into two modules. One module is used to determine, from the registered addresses, the addresses that can be used for amplification for the second data source, that is, the registered addresses equivalent to the addresses provided by the second data source. The second module can be used to perform multi-data source fusion based on the amplified situation to determine the confidence levels of the addresses provided or amplified by each data source belonging to the real addresses of the corresponding factories. Among them, in the process of fusion, the reliability factors of each data source as a whole can be considered to obtain a more accurate fusion result. The specific factory address verification results can be stored in the factory database. In this way, the factory address information can be provided to the "ground promotion" staff based on this database, so as to facilitate the "ground promotion" staff to conduct on-site visits to specific factories and other work.

[0066] The following details the specific implementation solutions provided by the embodiments of the present application.

[0067] First, the embodiments of the present application provide a method for verifying factory addresses. Referring to Figure 2 , the method specifically may include:

[0068] S201: Determine multiple factories to be verified for addresses, determine the first addresses corresponding to the multiple factories through a first data source, and collect second addresses for the multiple factories from multiple second data sources.

[0069] In specific implementation, first, information such as the names of multiple factories can be obtained, and then, address information of these factories can be collected. Specifically, first, the registered addresses of each factory can be collected. The so-called registered address refers to the address registered on the company's business license. Of course, in actual applications, this "registered address" may also have other names. Therefore, in the embodiments of the present application, it can be referred to as the first address. The characteristic of this first address is that as long as the name information of the factory is known and queried through a first data source such as a website provided by the industrial and commercial bureau and other departments, the corresponding registered address can basically be queried. That is, the coverage of this first data source for factories is relatively high. Of course, as described in the background art section, since the registered addresses of many factories may be inconsistent with the actual business addresses, only obtaining the first addresses of the factories is not enough, and more address information needs to be obtained from other data sources.

[0070] Among them, there can be various other data sources. For example, it can include the POI address data source provided by the map information system, or the factory tax information data source, or the merchant address data source provided in the relevant e-commerce system, or the factory address data source provided by the industrial park (for example, an industrial park may announce the factory occupancy situation in the park, etc.), or the address data source that the "ground promotion" operators have collected, and so on.

[0071] Here, it should be noted that the coverage of various second data sources for factories may be relatively low. That is to say, assuming that the second data sources are the above 5 types, however, for the same factory, only some of the data sources may be able to provide address information for this factory. For example, if a certain factory has not purchased the POI service in the map information system, then the POI data source of the map information system cannot provide address information for this factory, and so on. Therefore, in the process of specifically collecting the second address, it is collected according to the situation of the second address that the specific second data source can provide. That is, for a specific factory, if a data source can provide the address information of this factory, it can be collected as the second address. Otherwise, the address information provided by this data source for this factory is empty, and the address information can continue to be collected from other second data sources for this factory, and so on.

[0072] S202: For some or all of the multiple second data sources, determine the first address equivalent to the second address provided by the second data source from the first addresses respectively corresponding to the multiple factories.

[0073] After completing the address collection from various data sources for a factory set, in the embodiments of the present application, first, the address of the second data source can be amplified. Among them, since there can be various second data sources, therefore, the address of all the second data sources can be amplified respectively, or only the address of some of the data sources can be amplified. For example, for the latter, since the effectiveness of the POI address data source provided by the map information system is often relatively high, that is, the probability that the POI address belongs to the real address of the factory is relatively high, therefore, the address of this POI address data source can be amplified to increase the number of effective addresses participating in the multi-data source fusion calculation in the future, and so on.

[0074] Among them, when specifically amplifying the address of the second data source, a specific amplification process can be executed for the specific second data source respectively. For example, for one of the second data sources, first, sample data can be selected for this second data source from multiple factories (i.e., the total original number of factories), and positive samples and negative samples can be determined therefrom. Specifically, in one implementation manner, the first addresses of some factories that can provide the second address by the second data source can be determined as the sample data of this second data source. That is to say, assuming that the total number of factories to be verified for the address is m1, and a certain second data source can provide the second address for m2 of these factories, then these m2 factories can be used as the sample data of this second data source. After that, the distance between the first address and the second address of the same factory can be calculated. If the distance is less than the first target threshold, the first address of the corresponding factory is determined as a positive sample, otherwise as a negative sample.

[0075] That is to say, first, the text addresses of the first address and the second address can be converted into longitude and latitude by using APIs provided by relevant map information systems, etc., and then the distance between the first address and the second address of each factory is calculated. If the distance between the first address and the second address is less than 1 km (or other thresholds), the first addresses of these factories can be determined as the positive samples of the current second data source.

[0076] Among them, when specifically calculating the distance between the first address and the second address of the factory, it can be carried out in the following way: Given the longitude and latitude (j1, w1), (j2, w2) of two locations, after converting the angles into radians, the distance d between the two locations can be obtained through the spherical distance formula:

[0077]

[0078]

[0079] Among them, R is the radius of the earth.

[0080] In addition, as mentioned above, since specific factories may have the characteristic of being clustered and distributed within the industrial park, therefore, clustering processing can also be performed on the full amount of factories that need to be verified for the address specifically, and then by analyzing the specific clustering results, it is judged whether the conditions of the industrial park are met. If so, the registered addresses of the factories included therein may all have relatively high validity.

[0081] In addition, when performing clustering, the distances between the first addresses corresponding to multiple factories can be calculated pairwise, and then, each first address within the area where the first addresses are densely distributed can be determined as a cluster. Among them, when calculating the pairwise distances between multiple first addresses, the amount of data to be calculated is relatively large. Therefore, in specific implementation, the amount of calculation can be reduced in various ways.

[0082] For example, on the one hand, since a specific industrial park usually has the characteristic of being located within a certain target zoning range, that is, if two factories are located in the same industrial park, they are usually also located within the same target zoning range. Otherwise, if two factories already belong to different target zoning ranges, they usually do not belong to the same industrial park. Therefore, when performing clustering, the first addresses can also be first divided into multiple groups according to the target zoning ranges where the first addresses of each factory are located, and then the pairwise distances of the first addresses are calculated within each group. In this way, the number of pairwise combinations of the first addresses can be reduced, and the workload of calculating pairwise distances can be reduced.

[0083] In addition, when specifically calculating the pairwise distances of the first addresses, in addition to directly calculating the spherical distance based on the longitude and latitude between the first addresses of different factories, the method of local Euclidean coordinate approximation can also be used to calculate the distances between pairwise locations, so as to reduce the amount of calculation. Specifically, also based on a specific clustering algorithm running within the target zoning range of a certain unit, therefore, the boundary of a target zoning range on the sphere can be mapped to the boundary on a plane. For example, it can be a trapezoid on a plane, and then, the spherical longitude and latitude of the first addresses within the range are mapped into this trapezoid boundary, and the distances between pairwise locations are calculated based on the plane coordinates within the plane trapezoid. Through calculation verification, the error between the Euclidean distance and the spherical distance obtained after this method is mapped is within 1%.

[0084] After calculating the distances between multiple first addresses pairwise, the first addresses can be clustered based on methods such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise, density-based clustering algorithm). Different from the partitioning and hierarchical clustering methods, DBSCAN can define a cluster as the largest set of density-connected points, can divide the area with sufficient high density into clusters, and can discover clusters of any shape in a spatial database with noise. Specifically, given a distance function dist, a distance parameter d, and a core minimum sample number min_pts, the clusters found by this method can meet three conditions:

[0085] For each point p in each cluster C, there exists a point q ∈ C such that dist(p, q) ≤ d;

[0086] For any two clusters C and C′, min (p,q) ∈C×C′ dist(p, q) > d;

[0087] The size of each cluster C satisfies |C| > min_pts.

[0088] That is to say, by setting the above conditions, it can be ensured that the density of the clusters is relatively high. That is, for each point in the cluster, a point can be found such that the distance between the two is less than d. Moreover, there is no intersection between different clusters, and the number of points contained in a single cluster cannot be too small (otherwise it may not conform to the actual situation of the industrial park).

[0089] After obtaining multiple clusters through clustering, the first addresses within the obtained multiple clusters can be analyzed respectively to screen out the target clusters, and each first address in the target clusters can be determined as the first address equivalent to the second address provided by the second data source.

[0090] Among them, specifically when analyzing each cluster, for one of the clusters, the number of the equivalent first addresses contained in the cluster can be predicted based on the number of first addresses included in the cluster, the number of first addresses belonging to the sample data, and the number of positive samples therein. If the proportion of the number of the equivalent first addresses contained in the cluster exceeds the second target threshold, it is determined as the target cluster.

[0091] Specifically, when predicting the number of the equivalent first addresses contained in the cluster, a confidence estimation algorithm based on the hypergeometric distribution can be used to generate a function with the number of first addresses included in the cluster, the number of first addresses belonging to the sample data, and the number of positive samples therein as known parameters and the number of the equivalent first addresses contained in the cluster as an unknown parameter.

[0092] That is, taking the second data source as the POI address data source as an example, since the validity of the POI address is relatively high, therefore, for the task of analyzing the clusters, it is to estimate the proportion of valid addresses in a cluster in the tasks of identity authenticity and address validity (if the proportion of valid addresses in a cluster is relatively high, it can be proved that the probability of belonging to an industrial park is relatively high. Correspondingly, the first addresses of each factory in this cluster can be regarded as addresses equivalent to the POI address). However, based on the above tasks, no particularly suitable features have been determined at present. Therefore, a pure statistical method can be used to estimate the proportion of valid addresses in each cluster. That is, assuming that a cluster contains M factories, and the first addresses of m factories are sample addresses of the current second data source, and the number of positive samples among these sample addresses is k. Assuming that among the M factories, the number of addresses equivalent to this second data source is N, then it follows a hypergeometric distribution G(M, N, m), and:

[0093]

[0094] Among them, when M, m, and k are known, the posterior probability P r [N≥t,M,m,k] can be calculated using Bayes' formula:

[0095]

[0096] When the prior distribution of N takes a uniform distribution,

[0097]

[0098] In the embodiments of the present application, M, m, and k are known, and the unknown parameter N needs to be estimated. Specifically, in implementation, the estimation of the parameter N can be achieved in various ways. For example, in one way, maximum likelihood estimation can be used for estimation, and it can be proved that However, using maximum likelihood estimation often results in unreliable results. For example, assuming M = 50, m = 5, k = 5, although maximum likelihood estimation gives N = 50 (the point with the maximum probability in the figure as Figure 3 shown), the posterior point estimate at this place is only 0.12, and this value is on the low side.

[0099] Therefore, in the preferred embodiment of the present application, the method based on the confidence interval can be used to predict the number of equivalent first addresses. That is, the method of the lower bound of the confidence interval commonly used in risk control can be used to estimate

[0100]

[0101] where α is the significance parameter. Assuming α = 0.05 here, that is with a probability exceeding 95%. Using the above formula, a reliable estimate of N can be effectively obtained. If M = 50, m = 5, and k = 5, we can get corresponding to a confidence level of Figure 3 the area of the shaded part in

[0102] That is to say, when estimating by the maximum likelihood method, the goal is to maximize P r and the main focus is on Figure 3 the heights of the points in. By the method based on the confidence interval, the goal is that the probability that N is greater than or equal to t exceeds the requirement (e.g., 95%). Therefore, the main focus is on Figure 3 the area of the shaded part in. That is, by estimating N in this way, it can be ensured that the probability that the number of equivalent addresses is greater than the estimated value exceeds the preset goal, e.g., 95%, so as to obtain a batch of first addresses equivalent to the addresses provided by the current second data source.

[0103] It should be noted here that the validity of the addresses provided by various different second data sources may vary. Among them, the validity of the second addresses provided by the POI address data source is relatively high, that is, usually the real address of a specific factory. Therefore, in specific implementation, the equivalent addresses can be mainly obtained for this POI address data source. In this way, the first addresses specifically equivalent to the second addresses provided by the POI address data source can be the first addresses with the same confidence level as the POI address. Since the confidence level of this kind of equivalent address is also relatively high, the area corresponding to the specific cluster can be determined as an industrial park. At this time, this cluster can be determined as the target cluster, and then each first address in this target cluster can be determined as the first address equivalent to the POI address provided by the POI address data source. In this way, the number of valid addresses participating in the subsequent fusion calculation can be increased, which is beneficial to improving the probability of obtaining an accurate factory address.

[0104] S203: Use the equivalent first addresses to amplify the second addresses provided by the second data source.

[0105] After obtaining multiple equivalent first addresses corresponding to a specific second data source, the second addresses provided by the second data source can be amplified. That is, assume that a certain second data source could originally provide address information for n1 factories. Through the calculation in the previous step, n2 first addresses equivalent to the addresses provided by this second data source are determined from the first addresses. Then the number of addresses corresponding to this second data source can be amplified to n1 + n2. After that, based on the amplified addresses, fusion calculations between different data sources can be performed.

[0106] S204: Based on the address amplification results corresponding to each second data source, perform multi-data-source fusion processing to determine the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses.

[0107] Specifically, after obtaining the address amplification results corresponding to each second data source, multi-data-source fusion processing can be performed to determine the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses. Among them, specifically when performing multi-data-source fusion processing, an algorithm model for multi-data-source fusion can be constructed to predict the real address of a factory based on the addresses provided and / or amplified by multiple data sources for the same factory, and determine the confidence level of the address provided by the data source for the factory according to the distance between the address provided by the data source for a factory and the prediction result of the real address corresponding to the factory.

[0108] Among them, in order to improve the accuracy of fusion calculation, the specific algorithm model can include probability parameters for characterizing the reliability of the addresses provided by each data source; in this way, when predicting the real address of a factory, the algorithm model can be iterated multiple times, and the probability parameters can be updated during each iteration. After the algorithm converges, the output result of the corresponding algorithm model can be determined as the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses.

[0109] Specifically, the spherical distance between the longitude and latitude of the real address of a factory and the longitude and latitude of the address provided or amplified by the second data source can be comprehensively used to measure the accuracy of the address. Through analysis, it is found that the distribution of factory density generally follows a logarithmic distribution. Based on this characteristic, it can be assumed that the longitude and latitude provided by each data source j for factory i follows the following distribution:

[0110]

[0111] Among them, is the longitude and latitude of the real address, p j and λ j are the probability parameters corresponding to data source j, used to characterize the reliability of the address provided by data source j. Among them, if p j is relatively large, then the possibility of data source j generating an accurate address is small. On the contrary, if p j is relatively small, then the possibility of data source j generating an accurate address is large. If λ j is relatively large, then the possibility of data source j generating an accurate address is also relatively large. On the contrary, if p j is relatively small, then the possibility of data source j generating an accurate address is also relatively small.

[0112] Assume The latitude and longitude provided by data source j for factory i and the latitude and longitude of the true address of this factory i The distance between them is N i represents the number of second data sources that provide latitude and longitude (including the latitude and longitude of the amplified addresses) for the same factory. The log-likelihood function for the estimation of the true address can be:

[0113]

[0114] For the true address The estimation is transformed into a problem of maximizing the log-likelihood, maxLL. In specific implementation, LL can be optimized through an iterative algorithm. For example, the steps are as follows:

[0115] 1. Initialize the predicted latitude and longitude

[0116]

[0117] 2. Update the probability parameters λ j , p j :

[0118]

[0119]

[0120] 3. Fix the probability parameters and update the predicted latitude and longitude

[0121]

[0122] 4. Determine whether it converges. If it converges, output; if not, jump to step 2.

[0123] Through the above method, the confidence levels of the addresses provided / amplified by each data source can be obtained. For the same factory, the address with a higher confidence level among the addresses provided / amplified by each data source can be selected as the final fusion result. In this way, an address that may be the true address can be determined for a specific factory.

[0124] In summary, through the embodiments of the present application, for factories that require address verification, address information can be collected for specific factories from multiple data sources. This includes determining the first addresses corresponding to multiple factories through a first data source, and collecting second addresses for the multiple factories from multiple second data sources. Subsequently, for some or all of the multiple second data sources, a first address equivalent to the second address provided by the second data source can be determined from the first addresses corresponding to the multiple factories respectively, and the second address provided by the second data source can be amplified using the equivalent first address. Then, based on the address amplification results corresponding to each second data source, multi-data source fusion processing can be performed to determine the confidence levels that the addresses provided or amplified by the multiple data sources for each factory belong to real addresses. In this way, potential connections between different data sources can be discovered, enabling the first addresses provided by the first data source with a relatively high coverage to amplify the second addresses provided by the second data source, thereby enhancing the coverage of the second data source for factories. When performing multi-data source fusion processing on this basis, it is beneficial to obtain more accurate address information for more factories.

[0125] It should be noted that the embodiments of the present application may involve the use of user data. In practical applications, user-specific personal data can be used in the solutions described herein within the scope permitted by applicable laws and regulations, provided that the requirements of the applicable laws and regulations of the country where the user is located are met (for example, the user gives clear consent, and the user is effectively notified, etc.).

[0126] Corresponding to the foregoing method embodiments, the embodiments of the present application further provide a factory address verification device. Refer to Figure 4 , and the device may include:

[0127] An address collection unit 401, configured to determine multiple factories to be verified for addresses, determine the first addresses corresponding to the multiple factories through a first data source, and collect second addresses for the multiple factories from multiple second data sources;

[0128] An equivalent address determination unit 402, configured to, for some or all of the multiple second data sources, determine a first address equivalent to the second address provided by the second data source from the first addresses corresponding to the multiple factories respectively;

[0129] An address amplification unit 403, configured to amplify the second address provided by the second data source using the equivalent first address;

[0130] A multi-data source fusion unit 404, configured to perform multi-data source fusion processing based on the address amplification results corresponding to each second data source, to determine the confidence levels that the addresses provided or amplified by the multiple data sources for each factory belong to real addresses.

[0131] Among them, the equivalent address determination unit may specifically include:

[0132] A sample determination subunit, which is used to, for one of the second data sources, determine, from the multiple factories, the first addresses of the part of the factories that the second data source can provide with second addresses as the sample data of the second data source, and determine the distance between the first address and the second address of the same factory. If the distance is less than the first target threshold, the first address of the corresponding factory is determined as a positive sample, otherwise as a negative sample;

[0133] A clustering generation subunit, which is used to determine, according to the distances between the first addresses corresponding to the multiple factories pairwise, each first address within the region where the first addresses are densely distributed as a cluster;

[0134] A cluster analysis subunit, which is used to, by analyzing the first addresses within the obtained multiple clusters respectively, screen out the target cluster, and determine each first address within the target cluster as the first address equivalent to the second address provided by the second data source.

[0135] Among them, the clustering generation subunit may specifically be used to: within the same target zoning range, calculate the distances between the first addresses corresponding to the multiple factories pairwise, and determine each first address within the region where the first addresses are densely distributed within the same target zoning range as a cluster.

[0136] Specifically, the clustering generation subunit may be used to:

[0137] Map the boundary of the target zoning range on the spherical surface into the boundary on the plane, and approximately calculate the distances between the first addresses corresponding to the multiple factories pairwise through local Euclidean coordinates.

[0138] Among them, the cluster analysis subunit may specifically be used to:

[0139] For one of the clusters, based on the number of first addresses included in the cluster, the number of first addresses that belong to the sample data, and the number of positive samples among them, predict the number of the equivalent first addresses included in the cluster. If the proportion of the number of the equivalent first addresses included in the cluster exceeds the second target threshold, it is determined as the target cluster.

[0140] Specifically, a confidence estimation algorithm based on the hypergeometric distribution can be used to generate a function with the number of first addresses included in the cluster, the number of first addresses that belong to the sample data, and the number of positive samples among them as known parameters and the number of the equivalent first addresses included in the cluster as an unknown parameter, and based on the method of confidence intervals, predict the number of the equivalent first addresses.

[0141] Among them, the second data source includes a point of interest (POI) address data source provided by a map information system, and the second address includes a POI address;

[0142] The first address equivalent to the second address provided by the second data source includes:

[0143] A first address having the same confidence level as the POI address;

[0144] If the proportion of the number of the equivalent first addresses included in the cluster exceeds a second target threshold, it is determined that the area corresponding to the cluster is an industrial park, and the cluster is determined as the target cluster, and each first address in the target cluster is determined as a first address equivalent to the POI address provided by the POI address data source.

[0145] Among them, the multi-data source fusion unit can specifically be used for:

[0146] Construct an algorithm model for multi-data source fusion to predict the real address of a factory according to the addresses and / or amplified addresses provided by multiple data sources for the same factory, and determine the confidence level of the address provided by the data source for the factory according to the distance between the address provided by the data source for a certain factory and the prediction result of the real address corresponding to the factory.

[0147] Among them, the algorithm model includes probability parameters for characterizing the reliability of the addresses provided by each data source;

[0148] The multi-data source fusion unit can specifically be used for:

[0149] Perform multiple rounds of iteration on the algorithm model, and update the probability parameters during each iteration;

[0150] After the algorithm converges, determine the confidence level that the address provided or amplified by each data source for each factory belongs to the real address according to the output result of the corresponding algorithm model.

[0151] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the foregoing method embodiments are implemented.

[0152] And an electronic device, including:

[0153] One or more processors; and

[0154] A memory associated with the one or more processors, the memory being configured to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of the foregoing method embodiments.

[0155] Wherein, Figure 5 An exemplary architecture of an electronic device is shown, which may specifically include a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, and a memory 520. The above-mentioned processor 510, video display adapter 511, disk drive 512, input / output interface 513, network interface 514, and the memory 520 may be communicatively connected via a communication bus 530.

[0156] Wherein, the processor 510 may be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solution provided by the present application.

[0157] The memory 520 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 520 may store an operating system 521 for controlling the operation of the electronic device 500, and a basic input / output system (BIOS) for controlling the low-level operations of the electronic device 500. Additionally, a web browser 523, a data storage management system 524, a factory address verification processing system 525, etc. may also be stored. The above-mentioned factory address verification processing system 525 may be the application program that specifically implements the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 520 and are called and executed by the processor 510.

[0158] The input / output interface 513 is used to connect to an input / output module to achieve information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input devices may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices may include a display, a speaker, a vibrator, an indicator light, etc.

[0159] The network interface 514 is used to connect to a communication module (not shown in the figure) to enable communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0160] The bus 530 includes a path for transmitting information between various components of the device (such as the processor 510, video display adapter 511, disk drive 512, input / output interface 513, network interface 514, and memory 520).

[0161] It should be noted that although the above device only shows the processor 510, video display adapter 511, disk drive 512, input / output interface 513, network interface 514, memory 520, bus 530, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.

[0162] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0163] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0164] The above has introduced in detail the factory address verification method, device and electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A factory address verification method, characterized in that, Including: Determine multiple factories to be subject to address verification, determine the first addresses corresponding to the multiple factories through a first data source, and collect second addresses for the multiple factories from multiple second data sources; For some or all of the multiple second data sources, from the multiple factories, determine the first addresses of the partial factories that can provide second addresses by one of the second data sources as the sample data of the second data source, and determine the distance between the first address and the second address of the same factory. According to the distances between the first addresses corresponding to the multiple factories pairwise, determine each first address within the region where the first addresses are densely distributed as a cluster, and screen out the target cluster by analyzing the first addresses within the obtained multiple clusters respectively, so as to determine each first address within the target cluster as the first address equivalent to the second address provided by the second data source; Utilize the equivalent first addresses to amplify the second addresses provided by the second data source; Based on the address amplification results corresponding to each second data source, perform multi-data-source fusion processing to determine the confidence levels that the addresses provided or amplified by multiple data sources for each factory belong to real addresses.

2. The method according to claim 1, characterized in that It also includes: If the distance between the first address and the second address of the same factory is less than the first target threshold, determine the first address of the corresponding factory as a positive sample, otherwise as a negative sample.

3. The method according to claim 1, wherein the step of determining each first address within the region where the first addresses are densely distributed as a cluster according to the distances between the first addresses corresponding to the multiple factories pairwise includes: Within the same target zoning range, calculate the distances between the first addresses corresponding to the multiple factories pairwise, and determine each first address within the region where the first addresses are densely distributed within the same target zoning range as a cluster.

4. The method according to claim 3, wherein the step of calculating the distances between the first addresses corresponding to the multiple factories pairwise within the same target zoning range includes: Map the boundary of the target zoning range on the sphere into the boundary on the plane, and approximately calculate the distances between the first addresses corresponding to the multiple factories pairwise through local Euclidean coordinates.

5. The method according to claim 2, wherein the step of analyzing the first addresses within the obtained multiple clusters respectively includes: For one of the clusters, predict the number of the equivalent first addresses included in the cluster according to the number of the first addresses included in the cluster, the number of the first addresses that belong to the sample data, and the number of the positive samples therein. If the proportion of the number of the equivalent first addresses included in the cluster exceeds the second target threshold, determine it as the target cluster.

6. The method according to claim 5, wherein the step of predicting the number of the equivalent first addresses included in the cluster includes: A confidence estimation algorithm based on the hypergeometric distribution generates the number of first addresses included in a cluster, where the number of first addresses belonging to the sample data and the number of positive samples are known parameters, and the number of equivalent first addresses included in the cluster is a function of an unknown parameter. Based on the method of confidence intervals, the number of equivalent first addresses is predicted.

7. The method according to claim 5 or 6, characterized in that The second data source includes a point of interest (POI) address data source provided by a map information system, and the second address includes a POI address; The first address equivalent to the second address provided by the second data source includes: A first address having the same confidence level as the POI address; If the proportion of the number of equivalent first addresses included in the cluster exceeds a second target threshold, it is determined that the area corresponding to the cluster is an industrial park, and the cluster is determined as the target cluster, and each first address in the target cluster is determined as a first address equivalent to the POI address provided by the POI address data source.

8. The method according to claim 1, characterized in that The multi-data source fusion processing based on the address amplification results corresponding to each second data source includes: Constructing an algorithm model for multi-data source fusion to predict the real address of a factory according to the addresses and / or amplified addresses provided by multiple data sources for the same factory, and determining the confidence level of the address provided by the data source for the factory according to the distance between the address provided by the data source for a certain factory and the prediction result of the real address corresponding to the factory.

9. The method according to claim 8, characterized in that The algorithm model includes probability parameters for characterizing the reliability of the addresses provided by each data source; The prediction of the real address of the factory includes: Performing multiple rounds of iteration on the algorithm model and updating the probability parameters during each iteration; After the algorithm converges, determining the confidence level that the output results of the corresponding algorithm model indicate that the addresses provided or amplified by multiple data sources for each factory belong to the real address.

10. A factory address verification device, characterized in that, It includes: An address collection unit for determining multiple factories to be verified for addresses, determining the first addresses corresponding to the multiple factories through a first data source, and collecting second addresses for the multiple factories from multiple second data sources; An equivalent address determination unit for, for some or all of the multiple second data sources, determining, from the multiple factories, the first addresses of some factories that can provide second addresses by one of the second data sources as the sample data of the second data source, determining the distance between the first address and the second address of the same factory, determining each first address in the area where the first addresses are densely distributed as a cluster according to the distances between the first addresses corresponding to the multiple factories, and screening out the target cluster by analyzing the first addresses in the obtained multiple clusters respectively, so as to determine each first address in the target cluster as a first address equivalent to the second address provided by the second data source; An address amplification unit, configured to amplify the second address provided by the second data source by using the equivalent first address; A multi-data source fusion unit, configured to perform multi-data source fusion processing based on the address amplification results corresponding to each second data source, so as to determine the confidence levels of the addresses provided or amplified by multiple data sources for each factory being real addresses.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, Comprising: One or more processors; And A memory associated with the one or more processors, the memory being used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method according to any one of claims 1 to 9 are executed.

Citation Information

Patent Citations

  • Computing method based on distributed memory system and related device

    CN107770261A

  • Address Point Data Mining

    US20150031397A1