Multi-source interest point matching method and system based on multi-attribute feature similarity

By adopting a multi-attribute feature similarity fusion method in multi-source POI data matching, the problem of unsatisfactory POI matching accuracy in the prior art is solved, especially when processing Chinese data, and higher matching accuracy and applicability are achieved.

CN120030358APending Publication Date: 2025-05-23Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510068567.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The accuracy of the existing multi-source POI data matching methods is not ideal, especially when processing Chinese data, it is difficult for traditional string similarity calculation methods to accurately quantify the semantic similarity of attribute text.

Method used

A multi-source point of interest matching method based on multi-attribute feature similarity is adopted. By acquiring and preprocessing multi-source POI data, the name characters, character phonology, address characters, point of interest categories and contact similarity between target points of interest and candidate points of interest is calculated, and weighted fusion is performed to improve the accuracy of matching.

Benefits of technology

By comprehensively considering the similarity of multiple attribute features, the weak semantic constraints in traditional methods are broken, and the accuracy and applicability of POI matching is significantly improved, especially when processing Chinese data, the semantic similarity can be quantified more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030358A_ABST
    Figure CN120030358A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of geographic space data fusion, in particular to a multi-source interest point matching method and system based on multi-attribute feature similarity, and the method comprises the steps: obtaining multi-source POI data of a target region, carrying out the preprocessing of the multi-source POI data, obtaining the attribute features of a target interest point and a candidate interest point to be matched, and carrying out the matching of the target interest point and the candidate interest point to be matched; the attribute features comprise interest point name characters, interest point character phonics and shapes, interest point addresses, interest point categories and interest point contact information; and performing weighted fusion on the similarity between the attribute characteristics of the target interest point and the candidate interest point to be matched, and matching the similarity between the target interest point and the candidate interest point to be matched according to a weighted fusion result. According to the method, the contribution in the POI matching process can be more comprehensively improved, and the accuracy and applicability of POI matching are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geographic space data fusion, and in particular to a multi-source interest point matching method and system based on multi-attribute feature similarity. Background Art

[0002] POI, the full name is "Point of Interest". POI data can be understood as data information about specific points of interest, which can be geographical locations, commercial facilities, public facilities, transportation nodes and other types of information points that people may be interested in. POI data matching usually refers to the use of a specific similarity measurement method to determine whether two points of interest data from different data sources correspond to the same geographic entity in the physical world. Existing geographic entity matching methods mainly include three categories: methods based on spatial attributes, methods based on non-spatial attributes, and methods combining spatial and non-spatial attributes. Among them, the spatial-based matching method focuses more on the spatial position or geometric characteristics of the data. Generally speaking, the use of spatial position attributes is the most intuitive method to determine whether the geographic entity matches. However, the uncontrollable errors in the measurement process and the different encrypted coordinate systems used by map data service providers will cause coordinate offsets. When the data attribute content overlaps more, the use of a single attribute matching method will result in less accurate matching accuracy. The matching method combining spatial and non-spatial attributes realizes multi-source data matching by integrating multiple attribute elements. Therefore, many literatures tend to use matching methods based on the combination of spatial and non-spatial attributes.

[0003] At present, the matching method based on the combination of spatial attributes and non-spatial attributes has become increasingly mature. However, previous studies have focused more on how to reasonably allocate the weights of attributes. When calculating the similarity of attribute features, weak semantic constraints are still more often used, that is, the string similarity is directly calculated for attributes of various text types, resulting in different POIs with highly similar names or addresses being incorrectly matched as the same entity. In addition, compared with English words, Chinese lacks spaces as separators, and there are homophonic confusing words in Chinese. Therefore, directly using traditional string similarity calculation methods will result in the inability to accurately quantify the semantic similarity of the text in the attributes. Summary of the invention

[0004] To this end, the present invention provides a multi-source POI matching method and system based on multi-attribute feature similarity to solve the problem of unsatisfactory matching accuracy of existing multi-source POI data.

[0005] According to the design scheme provided by the present invention, on the one hand, a multi-source interest point matching method based on multi-attribute feature similarity is provided, comprising:

[0006] Acquire multi-source POI data of the target area, and pre-process the multi-source POI data to obtain attribute features of both the target POI and the candidate POI to be matched, wherein the attribute features include: POI name characters, POI character pronunciation and shape, POI address, POI category, and POI contact information;

[0007] The similarity between the attribute features of the target interest point and the candidate interest point to be matched is weightedly fused, and the similarity between the target interest point and the candidate interest point to be matched is matched according to the weighted fusion result.

[0008] As a multi-source POI matching method based on multi-attribute feature similarity of the present invention, further, the multi-source POI data is preprocessed, including:

[0009] Performing data cleaning on the multi-source POI data to obtain verified multi-source POI data, wherein the data cleaning includes deleting POI missing items and duplicate data by reviewing and verifying the data;

[0010] Each target point of interest in the verified POI data of one of the data sources is used as the center of a circle, and a buffer zone to be matched is generated with the specified search range as the radius;

[0011] The points to be matched in the buffer zone in the POI data of another data source are used as candidate interest points to be matched.

[0012] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, weighted fusion is performed on the similarity between the attribute features of the target interest point and the candidate interest point to be matched, including:

[0013] Calculate the name character similarity, character pronunciation and shape similarity, address character similarity, POI category similarity and contact information similarity between the target POI and the candidate POI;

[0014] The name character similarity, character pronunciation and shape similarity, address character similarity, POI category similarity and contact information similarity are weighted and fused according to the weight of each attribute parameter, so as to obtain the comprehensive similarity between the target POI and the candidate POI, wherein the weight of each attribute parameter is pre-set according to experience and / or experimental data.

[0015] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, calculating the name character similarity between the target interest point and the candidate interest point includes:

[0016] A regular expression is used to identify a name main body and an address additional name in the name of the point of interest, wherein the name main body is a name subject composed of a proper name and a business name, and the address additional name is used to additionally describe the address information of the point of interest;

[0017] The cosine similarity and Levenshtein distance similarity are used to calculate the similarity of the name body between the target interest point and the candidate interest point, and the Levenshtein distance similarity is used to calculate the similarity of the address additional name between the target interest point and the candidate interest point;

[0018] The name character similarity between the target POI and the candidate POI is obtained based on the name main body similarity and the address additional name similarity.

[0019] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, calculating the character sound and shape similarity between the target interest point and the candidate interest point includes:

[0020] The Dimsim algorithm is used to obtain the phonetic similarity between the name characters of the target interest point and the candidate interest points, and the glyph similarity between the names of the target interest point and the candidate interest points is calculated based on the Chinese character four-corner encoding library;

[0021] The pinyin similarity and the character shape similarity are added to obtain the character sound and shape similarity between the target interest point and the candidate interest points.

[0022] As a multi-source POI matching method based on multi-attribute feature similarity of the present invention, further, calculating the address character similarity between the target POI and the candidate POI comprises:

[0023] Extract address names at all levels from public data sets and construct an address segmentation corpus;

[0024] Use the address word library and regular expressions to segment the address location information of both the target point of interest and the candidate points of interest;

[0025] TF-IDF is used to calculate the label weight of each substring after word segmentation of the address location information of both the target interest point and the candidate interest point, and the address character similarity between the target interest point and the candidate interest point is obtained for each substring label address and label weight.

[0026] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, calculating the interest point category similarity between the target interest point and the candidate interest point includes:

[0027] Constructing semantic mapping relationships between POI category nodes from different data sources according to classification rules of POI data sources, wherein the semantic mapping relationships are used to describe semantic mapping of POI category codes from different POI data sources using a tree-like hierarchical structure;

[0028] The similarity of interest point categories between the target interest point and the candidate interest points is obtained based on the semantic mapping relationship between the interest point categories at each level in the tree hierarchy.

[0029] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, calculating the similarity of the connection mode between the target interest point and the candidate interest point includes:

[0030] It is determined whether there is a common substring in the contact information of the target interest point and the candidate interest point. If there is a common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 1. If there is no common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 0.

[0031] As a multi-source interest point matching method based on multi-attribute feature similarity of the present invention, further, weighted fusion is performed on the similarity between the attribute features of the target interest point and the candidate interest point to be matched, and further comprises:

[0032] Collect POI training data, and obtain comprehensive similarity between pairs of interest points based on name character similarity, character pronunciation and shape similarity, address similarity, category similarity, and contact information similarity between pairs of interest points in the POI training data;

[0033] Manually check the comprehensive similarity between interest point pairs in POI training and annotate binary matching labels to obtain a POI training sample set with annotated labels;

[0034] The support vector machine model is trained using the POI training sample set with annotated labels to obtain the weights of each attribute parameter based on the model training results.

[0035] In another aspect, the present invention further provides a multi-source interest point matching system based on multi-attribute feature similarity, comprising: a data acquisition module and a data matching module, wherein:

[0036] A data acquisition module is used to acquire multi-source POI data of a target area and pre-process the multi-source POI data to obtain attribute features of both the target POI and the candidate POI to be matched, wherein the attribute features include: POI name characters, POI character pronunciation and shape, POI address, POI category and POI contact information;

[0037] The data matching module is used to perform weighted fusion on the similarity between the attribute features of the target interest point and the candidate interest point to be matched, and match the similarity between the target interest point and the candidate interest point to be matched based on the weighted fusion result.

[0038] Beneficial effects of the present invention:

[0039] The present invention uses the similarity of multiple attributes including name, pronunciation, address, category and contact information to determine the matching of POIs, breaking the influence of existing weak semantic constraints on the calculation of attribute similarity, and can more comprehensively improve the contribution in the POI matching process, and improve the accuracy and applicability of POI matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a schematic diagram of a multi-source interest point matching process based on multi-attribute feature similarity in an embodiment;

[0041] Figure 2 This is a schematic diagram of the address attribute structure division in the embodiment;

[0042] Figure 3 This is a schematic diagram of the category mapping relationship in the embodiment;

[0043] Figure 4 It is a schematic diagram of the POI matching results of the experimental dataset in the embodiment. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.

[0045] In view of the problem that the accuracy of multi-source POI matching is not ideal, the present invention embodiment, see Figure 1 As shown, a multi-source interest point matching method based on multi-attribute feature similarity is provided, comprising:

[0046] S101, obtaining multi-source POI data of a target area, and preprocessing the multi-source POI data to obtain attribute features of both the target point of interest and the candidate points of interest to be matched, wherein the attribute features include: name characters of the point of interest, pronunciation and shape of the characters of the point of interest, address of the point of interest, category of the point of interest, and contact information of the point of interest.

[0047] Specifically, the multi-source POI data can be cleaned to obtain verified multi-source POI data, wherein the data cleaning includes deleting missing POI items and duplicate data by reviewing and verifying the data; taking each target interest point in the verified POI data from one of the data sources as the center of a circle and generating a buffer zone to be matched with a specified search range as the radius; and taking the points to be matched in the buffer zone in the POI data from another data source as candidate interest points to be matched.

[0048] By checking whether there is data missing or structural difference in the POI data, the data is cleaned. Data cleaning includes but is not limited to: removing POI items with missing addresses, deleting POI name symbols, capitalizing POI name letters, etc. Then, a buffer zone to be matched is generated with each point of the data as the center and the search range as the radius. The points to be matched in the buffer zone are used as the data to be matched at the point, and the reference data set is traversed to generate the corresponding buffer point group.

[0049] S102: weighted fusion is performed on the similarities between the attribute features of the target interest point and the candidate interest points to be matched, and the similarities between the target interest point and the candidate interest points to be matched are matched according to the weighted fusion result.

[0050] Specifically, the name character similarity, character pronunciation and shape similarity, address character similarity, interest point category similarity and contact information similarity between the target interest point and the candidate interest point can be calculated respectively; the name character similarity, character pronunciation and shape similarity, address character similarity, interest point category similarity and contact information similarity are weightedly fused according to the weight of each attribute parameter to obtain the comprehensive similarity between the target interest point and the candidate interest point, and the weight of each attribute parameter is pre-set based on experience and / or experimental data.

[0051] Calculating the name character similarity between the target interest point and the candidate interest points may include:

[0052] A regular expression is used to identify a name main body and an address additional name in the name of the point of interest, wherein the name main body is a name subject composed of a proper name and a business name, and the address additional name is used to additionally describe the address information of the point of interest;

[0053] The cosine similarity and Levenshtein distance similarity are used to calculate the similarity of the name body between the target interest point and the candidate interest point, and the Levenshtein distance similarity is used to calculate the similarity of the address additional name between the target interest point and the candidate interest point;

[0054] The name character similarity between the target POI and the candidate POI is obtained based on the name main body similarity and the address additional name similarity.

[0055] There is a certain pattern in the composition of POI Chinese names. The complete POI name can be decomposed into the following structure composed of several keywords:

[0056] POI name = administrative name + specific name + business name + organization name + general name + additional name (1)

[0057] In the map platform, most POI names can be basically described by the three parts of "proper name + business name + additional name". The first two items can be regarded as the main part of the name, and the third item is the additional description of the address. The rough address "additional name" has a relatively low overall similarity contribution in the overall similarity of the POI name due to its short length, but if it is ignored, it is very easy to cause incorrect matching of POI name attributes. Referring to the algorithm characteristics of cosine similarity and Levenshtein distance, in the embodiment of this case, two similarity weighted fusion calculation methods are used to obtain the name character similarity. According to the structural characteristics of the name attributes, regular expressions are used to identify the name body and the address additional name. The similarity weighted fusion algorithm is used to calculate the complex name body. The Levenshtein distance algorithm is used to calculate the address additional words in the POI name. The similarity results of the name body and the address additional words are weighted and summed to obtain the character similarity S of the POI name. name , various calculation formulas can be expressed as follows:

[0058]

[0059] S name =W sub (ω cos ·S cosine1 +ω LD ·S LD1 )+W suf ·S LD2 (5)

[0060] Among them, formula (2) is the cosine similarity calculation formula, formula (3) (4) is the Levenshtein distance similarity calculation formula, W sub , W suf are the weights of the name body and the address additional words, respectively. cosine1 , S LD1 , S LD2 They are the cosine similarity of the name body, the Levenshtein distance similarity of the name body, and the Levenshtein distance similarity of the address appended words.

[0061] Calculating the character sound and shape similarity between the target interest point and the candidate interest point may include:

[0062] The Dimsim algorithm is used to obtain the phonetic similarity between the name characters of the target interest point and the candidate interest points, and the glyph similarity between the names of the target interest point and the candidate interest points is calculated based on the Chinese character four-corner encoding library;

[0063] The pinyin similarity and the character shape similarity are added to obtain the character sound and shape similarity between the target interest point and the candidate interest points.

[0064] When matching two groups of words, the pronunciation and shape of Chinese characters are comprehensively considered. The Dimsim algorithm is used to compare the phonetic differences between the two characters to calculate the phonetic similarity, which is used to determine whether the pronunciation of the Chinese word is a homophonic word; secondly, a Chinese character four-corner encoding library is constructed to compare the digital encoding differences between Chinese characters to calculate the shape similarity. The two similarity results are weighted to obtain the comprehensive similarity score S of the two Chinese characters. ps , the calculation formula is as follows:

[0065] S ps =W pro ·Dis(u 1 ,u 2 )+W stru ·∑ k=0 δ(digit(A,k),digit(B,k)) (6)

[0066] Among them, W pro , W stru are the weights of pinyin similarity and glyph similarity, u 1 、u 2 They represent the phonetic codes of the characters respectively, and A and B represent the four-corner codes of the characters.

[0067] Calculating the address character similarity between the target POI and the candidate POI may include:

[0068] Extract address names at all levels from public data sets and construct an address segmentation corpus;

[0069] Use the address word library and regular expressions to segment the address location information of both the target point of interest and the candidate points of interest;

[0070] TF-IDF is used to calculate the label weight of each substring after word segmentation of the address location information of both the target interest point and the candidate interest point, and the address character similarity between the target interest point and the candidate interest point is obtained for each substring label address and label weight.

[0071] The address format can be summarized as X(province)X(city)X(district / county)X(street)X(road)X(building complex)X(building / building)X(unit)X(account number)X(others), and the closer to the end, the more specific the information. In this embodiment, the address names of various levels extracted from a large number of postal data sets can be used to construct a Chinese address segmentation corpus, such as Figure 2 As shown, the address vocabulary and regular expressions are used to segment the location information according to the format.

[0072] Although a custom word library is used, equal weights are used when calculating character strings. In this embodiment, a weighted address character similarity calculation method is used. The weighted method refers to the concept of word frequency-inverse text frequency (TF-IDF), adds corresponding tags to each substring after word segmentation, and calculates the weight value using TF-IDF. This solves the problem of lax attribute semantic constraints when processing address similarity to a high degree. TF-IDF formula and address character similarity S A The calculation formula is as follows:

[0073]

[0074]

[0075] TF-IDF=TF*IDF(9)

[0076] S A =∑ k=0 W k · S k (10)

[0077] Among them, formula (7), (8), (9) are the TF-IDF calculation formulas, formula (10) is the address character similarity calculation formula, W k Indicates the weight of each label address, S k Indicates the addresses of each label after word segmentation.

[0078] Calculating the similarity of interest point categories between the target interest point and the candidate interest points may include:

[0079] Constructing semantic mapping relationships between POI category nodes from different data sources according to classification rules of POI data sources, wherein the semantic mapping relationships are used to describe semantic mapping of POI category codes from different POI data sources using a tree-like hierarchical structure;

[0080] The similarity of interest point categories between the target interest point and the candidate interest points is obtained based on the semantic mapping relationship between the interest point categories at each level in the tree hierarchy.

[0081] In order to accurately and quickly distinguish various POIs and reduce data complexity, Amap and Tencent Map have established their own unique classification systems. In order to solve the problem of calculating the category similarity of POIs from the two data sources due to the differences in the classification systems, we follow the principles of scientificity, consistency, scalability and practicality when compiling the classification code system and build the following Figure 3 The semantic mapping relationship between category nodes of different classification systems is shown.

[0082] The category code consists of a three-level structure, which corresponds to three layers of tree nodes and three levels of classification codes. When calculating the category similarity of any two POIs, in this embodiment, whether each level of nodes satisfies the mapping relationship is used as the judgment standard, and a total of three similarity levels are sorted out: the first is that the two groups of POI categories have complete semantic mapping, in which case the category attribute similarity is 1; the second is that the two groups of POI categories only have parent nodes with semantic mapping relationships, in which case the category attribute similarity is 0-1; the third is that the two groups of POI categories do not have parent nodes with semantic mapping relationships, in which case the category attribute similarity is 0, and the category similarity S is 0. C The calculation formula is as follows:

[0083] D be =||u 1 u 2 ||=∑score i (11)

[0084]

[0085] Among them, u 1 with u 2 Represents two nodes in the category map N-ary tree.

[0086] Calculating the similarity of the contact information between the target interest point and the candidate interest points may include:

[0087] It is determined whether there is a common substring in the contact information of the target interest point and the candidate interest point. If there is a common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 1. If there is no common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 0.

[0088] A POI can have multiple contact information. When calculating the similarity of contact information, whether there is a common substring can be used as the basis for judging the similarity of contact information. When the contact information of the reference dataset and the matching candidate dataset have a common substring, the contact similarity is determined to be 1, indicating that the two POIs are highly matched in terms of contact similarity. Contact similarity S T The calculation formula is as follows:

[0089]

[0090] Among them, C gd Represents the contact attribute string of Amap POI, C tx A string representing the contact attribute of Tencent POI.

[0091] The weighted fusion of the similarity between the attribute features of the target interest point and the candidate interest point to be matched also includes:

[0092] Collect POI training data, and obtain comprehensive similarity between pairs of interest points based on name character similarity, character pronunciation and shape similarity, address similarity, category similarity, and contact information similarity between pairs of interest points in the POI training data;

[0093] Manually check the comprehensive similarity between interest point pairs in POI training and annotate binary matching labels to obtain a POI training sample set with annotated labels;

[0094] The support vector machine model is trained using the POI training sample set with annotated labels to obtain the weights of each attribute parameter based on the model training results.

[0095] The rationality of feature weight setting is related to the accuracy of matching results. In order to avoid the situation of "strengthening" or "weakening" certain features due to lack of human experience, the support vector machine binary classification algorithm is used in the embodiment of this case to reduce classification errors as much as possible. The weights of each attribute feature are generated by training the SVC model with manual proofreading data with binary matching labels, attribute features and comprehensive similarity calculation results, and a comprehensive similarity discrimination model is obtained for subsequent matching training of reference data sets. The calculation formula of the calculated comprehensive similarity Sim is as follows:

[0096] Sim=W N ·(ω 1 ·S N +ω 2 ·S ps )+W A ·S A +W C ·S c +W T ·S T (14)

[0097] Among them, each item W and ω represents the weight of the corresponding parameter item, and each item S represents the similarity of the corresponding attribute.

[0098] Furthermore, based on the above method, an embodiment of the present invention also provides a multi-source interest point matching system based on multi-attribute feature similarity, comprising: a data acquisition module and a data matching module, wherein:

[0099] A data acquisition module is used to acquire multi-source POI data of a target area and pre-process the multi-source POI data to obtain attribute features of both the target POI and the candidate POI to be matched, wherein the attribute features include: POI name characters, POI character pronunciation and shape, POI address, POI category and POI contact information;

[0100] The data matching module is used to perform weighted fusion on the similarity between the attribute features of the target interest point and the candidate interest point to be matched, and match the similarity between the target interest point and the candidate interest point to be matched based on the weighted fusion result.

[0101] In order to verify the effectiveness of this solution, the following is a further explanation based on experimental data:

[0102] Three areas of different sizes in Wuhou District of Amap (a total of 5040 POIs) were randomly selected as reference data, and the Tencent map POIs in the buffer zone were used as matching candidate data sets. The three processed data sets were matched using the trained comprehensive similarity discrimination model, and the accuracy, recall rate and F1 score of each reference data set were calculated to determine whether the matching effect meets the requirements. Table 1 and Figure 4 This paper presents the results and evaluation of the multi-source interest point matching calculation method based on multi-feature attribute similarity in this case.

[0103] Table 1 Evaluation matching results

[0104]

[0105] From Table 1 and Figure 4 It can be seen that the mentioned method has achieved good performance on the three test data sets. A total of 3755 geographic entities with matching points on the Amap and Tencent map platforms were identified, of which 35 were mismatched. The accuracy of correct matching among the matched data points in the data set has reached more than 98%, and the recall rates have reached 94.9%, 96.9% and 95.7% respectively, basically maintaining above 95%, indicating that the method proposed in this study still has an accurate description when judging whether multi-source POIs correspond to the same geographic entity, and is capable of correctly matching POIs to real entities. The F1 score, as the harmonic mean of precision and recall, is an intuitive measure of the performance of the classification model, and is also as high as more than 97% on the three test sets. This shows that the proposed multi-source point of interest matching algorithm model based on multi-attribute features has certain advantages in POI matching.

[0106] Unless otherwise specifically stated, the relative steps, numerical expressions and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0107] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0108] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0109] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.

[0110] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multi-source interest point matching method based on multi-attribute feature similarity, characterized in that: Include: Acquire multi-source POI data of the target area, and pre-process the multi-source POI data to obtain attribute features of both the target POI and the candidate POI to be matched, wherein the attribute features include: POI name characters, POI character pronunciation and shape, POI address, POI category, and POI contact information; The similarity between the attribute features of the target interest point and the candidate interest point to be matched is weightedly fused, and the similarity between the target interest point and the candidate interest point to be matched is matched according to the weighted fusion result.

2. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 1 is characterized in that: Preprocess multi-source POI data, including: Performing data cleaning on the multi-source POI data to obtain verified multi-source POI data, wherein the data cleaning includes deleting POI missing items and duplicate data by reviewing and verifying the data; Each target point of interest in the verified POI data of one of the data sources is used as the center of a circle, and a buffer zone to be matched is generated with the specified search range as the radius; The points to be matched in the buffer zone in the POI data of another data source are used as candidate interest points to be matched.

3. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 1 is characterized in that: The similarity between the attribute features of the target interest point and the candidate interest point to be matched is weighted and fused, including: Calculate the name character similarity, character pronunciation and shape similarity, address character similarity, POI category similarity and contact information similarity between the target POI and the candidate POI; The name character similarity, character pronunciation and shape similarity, address character similarity, POI category similarity and contact information similarity are weighted and fused according to the weight of each attribute parameter, so as to obtain the comprehensive similarity between the target POI and the candidate POI, wherein the weight of each attribute parameter is pre-set according to experience and / or experimental data.

4. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 3 is characterized in that: Calculate the name character similarity between the target interest point and the candidate interest points, including: A regular expression is used to identify a name main body and an address additional name in the name of the point of interest, wherein the name main body is a name subject composed of a proper name and a business name, and the address additional name is used to additionally describe the address information of the point of interest; The cosine similarity and Levenshtein distance similarity are used to calculate the similarity of the name body between the target interest point and the candidate interest point, and the Levenshtein distance similarity is used to calculate the similarity of the address additional name between the target interest point and the candidate interest point; The name character similarity between the target POI and the candidate POI is obtained based on the name main body similarity and the address additional name similarity.

5. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 3 is characterized in that: Calculate the character sound and shape similarity between the target interest point and the candidate interest points, including: The Dimsim algorithm is used to obtain the phonetic similarity between the name characters of the target interest point and the candidate interest points, and the glyph similarity between the names of the target interest point and the candidate interest points is calculated based on the Chinese character four-corner encoding library; The pinyin similarity and the character shape similarity are added to obtain the character sound and shape similarity between the target interest point and the candidate interest points.

6. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 3 is characterized in that: Calculate the address character similarity between the target POI and the candidate POI, including: Extract address names at all levels from public data sets and construct an address segmentation corpus; Use the address word library and regular expressions to segment the address location information of both the target point of interest and the candidate points of interest; TF-IDF is used to calculate the label weight of each substring after word segmentation of the address location information of both the target interest point and the candidate interest point, and the address character similarity between the target interest point and the candidate interest point is obtained for each substring label address and label weight.

7. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 3 is characterized in that: Calculate the similarity of interest point categories between the target interest point and the candidate interest points, including: Constructing semantic mapping relationships between POI category nodes from different data sources according to classification rules of POI data sources, wherein the semantic mapping relationships are used to describe semantic mapping of POI category codes from different POI data sources using a tree-like hierarchical structure; The similarity of interest point categories between the target interest point and the candidate interest points is obtained based on the semantic mapping relationship between the interest point categories at each level in the tree hierarchy.

8. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 3 is characterized in that: Calculate the similarity of the contact information between the target interest point and the candidate interest points, including: It is determined whether there is a common substring in the contact information of the target interest point and the candidate interest point. If there is a common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 1. If there is no common substring, the similarity of the contact information between the target interest point and the candidate interest point is determined to be 0.

9. The multi-source interest point matching method based on multi-attribute feature similarity according to claim 1 or 3, characterized in that: The similarity between the attribute features of the target interest point and the candidate interest point to be matched is weightedly fused, and also includes: Collect POI training data, and obtain comprehensive similarity between pairs of interest points based on name character similarity, character pronunciation and shape similarity, address similarity, category similarity, and contact information similarity between pairs of interest points in the POI training data; Manually check the comprehensive similarity between interest point pairs in POI training and annotate binary matching labels to obtain a POI training sample set with annotated labels; The support vector machine model is trained using the POI training sample set with annotated labels to obtain the weights of each attribute parameter based on the model training results.

10. A multi-source interest point matching system based on multi-attribute feature similarity, characterized in that: Contains: data acquisition module and data matching module, among which, A data acquisition module is used to acquire multi-source POI data of a target area and pre-process the multi-source POI data to obtain attribute features of both the target POI and the candidate POI to be matched, wherein the attribute features include: POI name characters, POI character pronunciation and shape, POI address, POI category and POI contact information; The data matching module is used to perform weighted fusion on the similarity between the attribute features of the target interest point and the candidate interest point to be matched, and match the similarity between the target interest point and the candidate interest point to be matched based on the weighted fusion result.