Map interest point supplementing method and device

By correcting and structural analysis of the address information entered by the user, extracting key elements and matching them with preset interest points, the problems of low replenishment efficiency and low matching success rate in the prior art are solved, and efficient and accurate supplementation of map interest points information is achieved.

CN120123447APending Publication Date: 2025-06-10SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510259081.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency and high cost in replenishing map interest points, and the address information processing of user feedback lacks an effective information processing mechanism, resulting in problems of spelling errors and irregular expressions, and the matching success rate is low.

Method used

By obtaining the query request input by the user, the address information is corrected and structured analysis is performed, key elements such as regional words, core words, category words and logical sub-points are extracted, and they are matched with the feature information of the preset interest points. If the match is successful, the address information is supplemented into the map as a new interest point.

Benefits of technology

It improves the success rate of interest point matching, realizes efficient and accurate supplementation of map interest point information, and meets users' needs for timeliness and richness of map information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123447A_ABST
    Figure CN120123447A_ABST
Patent Text Reader

Abstract

The invention discloses a map interest point supplementing method and device, and the method comprises the steps: obtaining a query request inputted by a target user, and the query request comprises address information; performing error correction processing on the address information; the address information subjected to error correction processing is subjected to structure analysis, to-be-matched elements are obtained, and the to-be-matched elements comprise at least one of region words, core words, category words and logic sub-points; matching the to-be-matched element with feature information of a preset point of interest; and if the matching is successful, taking the address information as a new interest point, and supplementing the new interest point to the target map. According to the method, the problems of error correction and deep analysis of the address information input by the user can be effectively solved, the interest point matching success rate is improved, map interest point information is more efficiently and accurately supplemented, and the requirements of the user for the accuracy and richness of map information are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maps, and particularly to a method and device for supplementing map points of interest. Background Art

[0002] In the fields of geographic information systems (GIS) and map applications, accurate and rich point of interest (POI) information is crucial for enhancing the user experience. Currently, the supplementation of map points of interest mainly relies on manual collection and entry, as well as a small amount of user feedback. The manual collection and entry method usually involves professional teams collecting and organizing information on various locations according to established standards. Although it ensures the accuracy of the data to a certain extent, it has the problems of low efficiency and high cost, and it is difficult to quickly respond to a large number of newly added locations and user personalized needs.

[0003] For the method of supplementing points of interest based on user feedback, although some new locations discovered by users can be obtained, due to the lack of an effective information processing mechanism, there are many problems. On the one hand, the address information input by users may have spelling mistakes, non-standard expressions, etc., making it difficult to accurately identify and process. On the other hand, for the address information input by users, the existing processing methods often simply match it with the existing point of interest information, lacking in-depth structured analysis of the address information, and unable to fully excavate the key elements in the address information, resulting in a low matching success rate, and a lot of valuable new point of interest information cannot be supplemented to the map in a timely manner. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and device for supplementing map points of interest for the above technical problems to solve at least one of the problems existing in the above prior art.

[0005] In a first aspect, a method for supplementing map points of interest is provided, including:

[0006] Obtaining a query request input by a target user, where the query request includes address information;

[0007] Performing error correction processing on the address information;

[0008] Performing structural analysis on the address information after error correction processing to obtain elements to be matched, where the elements to be matched include at least one of a geographical word, a core word, a category word, and a logical sub-point;

[0009] Matching the elements to be matched with the feature information of preset points of interest;

[0010] If the match is successful, taking the address information as a new point of interest and supplementing it to the target map.

[0011] In an embodiment, the performing error correction processing on the address information includes:

[0012] Correct the glyphs of each character in the address information;

[0013] Correct the pronunciation of each character in the address information.

[0014] In one embodiment, the correcting the glyphs of each character in the address information includes:

[0015] Match the address information with the preset point of interest character by character;

[0016] If the character in the address information fails to match the character in the preset point of interest, determine whether the characters before and after the failed matching character match successfully;

[0017] If the match is successful, determine the similarity between the address information and the failed matching character in the preset point of interest;

[0018] If the similarity is greater than the preset similarity threshold, modify the failed matching character.

[0019] In one embodiment, the determining the similarity between the address information and the failed matching character in the preset point of interest includes:

[0020] Determine the Cangjie code similarity between the address information and the failed matching character in the preset point of interest;

[0021] Determine the four-corner code similarity between the address information and the failed matching character in the preset point of interest;

[0022] Determine the stroke number similarity between the address information and the failed matching character in the preset point of interest;

[0023] Based on the Cangjie code similarity, four-corner code similarity, and stroke number similarity, determine the comprehensive similarity.

[0024] In one embodiment, the correcting the pronunciation of each character in the address information includes:

[0025] Convert each character in the address information into pinyin to obtain a pinyin sequence;

[0026] Determine whether there are a preset number of repeated pinyin combinations in the pinyin sequence;

[0027] If so, retain multiple Chinese characters corresponding to the repeated pinyin combination as one Chinese character.

[0028] In one embodiment, the element to be matched includes geographical information, and the matching the element to be matched with the feature information of the preset point of interest includes:

[0029] Query the preset geographical relationship dictionary to determine whether there is a parent-child relationship between the geographical information and the geographical information of the preset point of interest;

[0030] If there is a parent-child relationship, then the match is consistent.

[0031] In one embodiment, the element to be matched includes logical sub-points. The matching of the element to be matched with the feature information of the preset point of interest includes:

[0032] Match the logical sub-points with the logical sub-points of the preset point of interest;

[0033] If a continuous preset number of characters in the logical sub-points match successfully, then the match is consistent.

[0034] In one embodiment, the element to be matched includes a core word. The matching of the element to be matched with the feature information of the preset point of interest includes:

[0035] Construct an inverse document frequency dictionary, which includes inverse document frequency values corresponding to several points of interest;

[0036] Based on the inverse document frequency dictionary, calculate the term frequency-inverse document frequency scores corresponding to the core word and the core word in the preset point of interest respectively;

[0037] Determine the similarity between the term frequency-inverse document frequency score corresponding to the core word and the term frequency-inverse document frequency score corresponding to the core word in the preset point of interest;

[0038] When the similarity is greater than the preset similarity threshold, then the match is consistent.

[0039] In one embodiment, the element to be matched includes a category word. The matching of the element to be matched with the feature information of the preset point of interest includes:

[0040] Calculate the term frequency-inverse document frequency scores of the category word and the category word of the preset point of interest respectively;

[0041] Based on the term frequency-inverse document frequency scores, match the category word with the category word of the preset point of interest;

[0042] Determine the fingerprint codes corresponding to the category words with matching failures respectively;

[0043] Determine the similarity of the fingerprint codes;

[0044] If the similarity is greater than the preset similarity threshold, modify the term frequency-inverse document frequency scores corresponding to the category words with matching failures.

[0045] Second aspect, a map point of interest supplementing device is provided, including:

[0046] A query request obtaining unit, configured to obtain a query request input by a target user, where the query statement includes address information;

[0047] An error correction processing unit, configured to perform error correction processing on the address information;

[0048] A to-be-matched element obtaining unit, configured to perform structure parsing on the address information after error correction processing to obtain to-be-matched elements;

[0049] A matching unit, configured to match the to-be-matched elements with feature information of preset points of interest;

[0050] A point of interest supplementing unit, configured to, if the matching is successful, use the address information as a new point of interest and supplement it to the target map.

[0051] The above map point of interest supplementing method and device, the implementation of the method includes: obtaining a query request input by a target user, where the query request includes address information; performing error correction processing on the address information; performing structure parsing on the address information after error correction processing to obtain to-be-matched elements, where the to-be-matched elements include at least one of a regional word, a core word, a category word, and a logical sub-point; matching the to-be-matched elements with feature information of preset points of interest; if the matching is successful, using the address information as a new point of interest and supplementing it to the target map. This application can effectively solve the problems of error correction and in-depth parsing of address information input by users, improve the success rate of point of interest matching, and thus more efficiently and accurately supplement map point of interest information to meet the user's requirements for the accuracy and richness of map information. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 is a flowchart of a map point of interest supplementing method in an embodiment of the present invention;

[0054] Figure 2 is a structural diagram of a map point of interest supplementing device in an embodiment of the present invention;

[0055] Figure 3 is a schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] In one embodiment, as Figure 1 shown, a method for supplementing map points of interest is provided, including the following steps:

[0058] In step S110, obtain a query request input by a target user, where the query request includes address information;

[0059] Optionally, when the target user uses a map-related application or service, such as a map application, the user can enter a query request in the input box of the map application. The map application will monitor the content change in the input box in real time. When the user completes the input and triggers the search operation, the system will capture the text content in the input box and use it as the query request for subsequent processing. The query request may include address information, such as No. XX, XX Avenue, XX District.

[0060] It can be understood that the query request can also be entered by voice input. By converting the voice into text, the address information can be obtained. Or the map service provider can cooperate with other third-party applications or platforms to obtain relevant query requests input by users on these platforms through interface docking. Or a user feedback form can be set on the map application or related website, and users can submit newly discovered points of interest or relevant query requests by filling out the form. The form usually includes an address information input field, a point of interest type description field, etc. When the user fills out the form and submits it, the system will collect the data in the form and use it as the query request for processing. To encourage users to actively provide feedback, some reward mechanisms, such as points, coupons, etc., can also be set. Or when the positioning function of the user's device is enabled, the map application can automatically obtain possible query requests of the user according to the current geographical location information of the device.

[0061] In step S120, perform error correction processing on the address information;

[0062] Optionally, since the query request input by the user may contain errata, the glyph and pronunciation corrections may be performed before matching. Taking glyph correction as an example, the address information may be corrected character by character. If the characters before and after the character match the characters of the preset point of interest (POI), and the character does not match, the similarity between the character and the corresponding character in the preset point of interest may be calculated. If the similarity is greater than a preset similarity threshold, such as 90%, the character may be modified. Taking pronunciation as an example, the address information may be converted from text to pinyin, and the pinyin may be detected to determine if there are preset repeated pinyin combinations, such as haohaohao, which indicates that the input is incorrect. At this time, the Chinese character corresponding to the pinyin may be modified to the same character, such as "好好好" to "好", to remove repeated Chinese characters.

[0063] In step S130, the address information after error correction is structurally parsed to obtain elements to be matched, wherein the elements to be matched include at least one of a regional word, a core word, a category word, and a logical sub-point;

[0064] Optionally, the address information after error correction can be cleaned to remove special symbols, punctuation marks and other noise data that may interfere with the model judgment. Then the text is split by characters to obtain a character sequence, and feature extraction is performed for each character. The extracted features include but are not limited to the character itself, the context information of the character (such as the previous character and the next character of the current character), the character's part of speech (if the part of speech has been tagged), the position of the character in the text (such as whether it is the beginning, the end, etc.). The extracted features are combined into a feature vector, each character corresponds to a feature vector, and the extracted feature vector sequence is input into the name structure model CRF. The CRF model outputs the corresponding labels of each character in the address information. For example, "Beijing Haidian District Zhongguancun", "Beijing", "Hai", "Dian" and "District" belong to the region (administrative district), and "Zhong", "Guan" and "Cun" belong to the core words.

[0065] It should be noted that the name structuring model may include an input layer, a feature extraction layer, and a CRF layer. The input layer is used to receive the preprocessed character sequence, convert each character into a corresponding feature vector, and then input it into the model. For example, an Embedding Layer can be used to map characters to a low-dimensional vector space to capture the semantic information of the characters and reduce the data dimension at the same time. The feature extraction layer can use a Recurrent Neural Network (RNN) and its variants (such as LSTM, GRU) to effectively process the character sequence data and capture the context dependencies between characters. Through multiple layers of RNN networks, higher-level feature representations can be continuously extracted. After passing through the feature extraction layer, the extracted feature vector sequence is input into the CRF layer. The CRF layer calculates the label probability of each character based on the input feature vectors and considers the label dependencies of the entire sequence, and finally outputs the optimal label sequence.

[0066] Optionally, the training process of the CRF model is as follows: A large amount of POI (Point of Interest) place name data, such as 20,000 pieces, can be collected. These data cover various types of points of interest, such as shopping malls, restaurants, scenic spots, etc., and contain rich address information. Each piece of POI place name data is annotated character by character. The annotation categories include whether it belongs to a region (provincial and municipal administrative regions), core words, category words, and logical sub-points. The annotated data is divided into a training set, a validation set, and a test set. Usually, it is divided according to a certain ratio. Then, the training set data is input into the CRF model for iterative training. During the training process, the model calculates the probability of each character belonging to different categories based on the input annotated data and the designed feature functions. And an optimization algorithm (such as the gradient descent algorithm) can be used to adjust the parameters of the model to minimize the difference between the prediction result of the model on the training set and the annotated result. During the training process, the validation set is regularly used to evaluate the performance of the model. By calculating indicators such as accuracy, recall rate, and F1 value on the validation set, it is observed whether the model has overfitting or underfitting. According to the evaluation results on the validation set, the model is optimized and adjusted. Finally, the test set can be used to perform a final evaluation on the trained model. The trained and evaluated CRF model can be used for name structuring prediction of new address information.

[0067] In step S140, the element to be matched is matched with the feature information of the preset point of interest;

[0068] Optionally, both the element to be matched and the feature information may include at least one of a regional term, a core term, a category term, and a logical sub-point. Therefore, the regional term, core term, category term, and logical sub-point output by the name structuring model can be respectively matched with the corresponding feature information of the preset point of interest. For example, the regional term is matched with the regional term of the preset point of interest, the core term is matched with the core term of the preset point of interest, the category term is matched with the category term of the preset point of interest, and the logical sub-point is matched with the logical sub-point of the preset point of interest.

[0069] In step S150, if the match is successful, the address information is used as a new point of interest and added to the target map.

[0070] Optionally, if the address information matches the preset point of interest successfully, the address information is used as a new point of interest and added to the target map. Exemplarily, the address information can be arranged in a standard address format to clarify specific information such as province, city, district, street, and house number. And it is formatted according to the data format requirements of the target map. Connect to the database of the target map. And according to the table structure of the map database, the address information is inserted into the corresponding fields to generate a new point of interest, and a unique identifier (such as an ID number) is generated for the new point of interest. After adding the new point of interest, the index of the map is updated so that it can contain the information of the new point of interest, and the relevant cache data is updated. When the user accesses the map next time, according to the visualization rules and styles of the map, the new point of interest is displayed on the map in an appropriate form such as an icon and a label.

[0071] The embodiment of the present application provides a method and device for supplementing points of interest on a map, and the method is implemented by: obtaining a query request input by a target user, the query request including address information; performing error correction processing on the address information; performing structural analysis on the address information after error correction processing to obtain elements to be matched, the elements to be matched including at least one of regional words, core words, category words and logical sub-points; matching the elements to be matched with the feature information of a preset point of interest; if the match is successful, the address information is used as a new point of interest and supplemented to the target map. In the present application, by performing error correction processing on the address information, it is possible to effectively identify and correct spelling errors, non-standard expressions and other problems input by the user, so that subsequent information processing is based on accurate data, greatly improving the availability of address information. Using the name structured model to perform structural analysis on the address information, it is possible to accurately extract key elements such as regional words, core words, category words and logical sub-points, comprehensively and deeply understand the connotation of the address information, and provide richer and more valuable information for subsequent matching with the feature information of the preset point of interest. Based on the matching of the elements to be matched obtained by deep analysis and the characteristic information of the preset points of interest, compared with the traditional simple matching method, it can judge the relevance of address information and existing points of interest from multiple dimensions and in a more detailed manner, thereby significantly improving the matching success rate and allowing more new points of interest that meet the conditions to be discovered and supplemented. The entire process realizes the automated processing from user input to the supplement of points of interest, greatly reducing manual intervention, saving time and labor costs, and being able to quickly respond to query requests input by a large number of users, and timely add new points of interest to the target map, so that map information can be updated and improved more quickly, meeting users' needs for the timeliness and richness of map information.

[0072] In an embodiment of the present application, the error correction processing of the address information includes:

[0073] Correcting the glyphs of each character of the address information;

[0074] Correct the pronunciation of each character in the address information.

[0075] Optionally, the address information is corrected character by character. If the characters before and after the character match the characters of the preset point of interest (POI), but the character does not match, the similarity between the character and the corresponding character in the preset point of interest can be calculated. If the similarity is greater than a preset similarity threshold, such as 90%, the character can be modified. The address information can also be converted from text to pinyin, and the pinyin is detected. If it is determined that there are preset repeated pinyin combinations, such as haohaohao, it means that the input is incorrect. At this time, the Chinese character corresponding to the pinyin can be modified to the same character, for example, from "好好好" to "好", so as to remove repeated Chinese characters.

[0076] In an embodiment of the present application, correcting the glyphs of each character in the address information includes:

[0077] Matching the address information with the preset point of interest character by character;

[0078] If the character in the address information fails to match the character in the preset point of interest, determine whether the characters before and after the character with the matching failure match successfully;

[0079] If the match is successful, determine the similarity between the address information and the character with the matching failure in the preset point of interest;

[0080] If the similarity is greater than the preset similarity threshold, modify the character with the matching failure.

[0081] Optionally, the address information can be compared with the preset point of interest information character by character in sequence. If there is a character comparison failure, that is, the characters are inconsistent, it can be determined whether the characters before and after the character match. For example, if the address information is ABC and the preset point of interest information is ADC, the character B is inconsistent with D, but the characters AC before and after B are both consistent with the characters AC before and after D. Then the similarity between the characters B and D can be calculated. If the similarity is greater than the preset similarity threshold, for example, greater than 95%, then the character B needs to be modified, that is, B is modified to D.

[0082] In an embodiment of the present application, determining the similarity between the address information and the character with the matching failure in the preset point of interest includes:

[0083] Determine the similarity of the Cangjie codes between the address information and the character with the matching failure in the preset point of interest;

[0084] Determine the similarity of the four-corner codes between the address information and the character with the matching failure in the preset point of interest;

[0085] Determine the similarity of the stroke counts between the address information and the character with the matching failure in the preset point of interest;

[0086] Based on the similarity of the Cangjie codes, the similarity of the four-corner codes, and the similarity of the stroke counts, determine the comprehensive similarity.

[0087] The Cangjie code is a Chinese character input method encoding. For the characters with matching failures in the address information and the preset point of interest, obtain their Cangjie codes respectively. This can be achieved by querying the Cangjie code table or using the corresponding programming libraries. For example, in Python, some character encoding processing libraries can be used to obtain the corresponding Cangjie code string according to the character. And a suitable string similarity algorithm, such as the edit distance algorithm (Levenshtein distance), can be used to calculate the similarity between the two Cangjie code strings.

[0088] Four-corner coding is a method of encoding Chinese characters based on the strokes of the four corners of the characters. For characters that fail to match, their four-corner codes are obtained according to the four-corner coding rules. It can also be automatically obtained by looking up the four-corner coding table or using programming. For example, a mapping relationship of four-corner codes is established in the program, and the corresponding four-corner codes are output after the characters are input. And a suitable similarity algorithm, such as the edit distance algorithm (Levenshtein distance), can be used to calculate the similarity between two four-corner codes.

[0089] For characters that fail to match, the number of strokes of each character can be determined according to the stroke rules of Chinese characters. The difference in the number of strokes of two characters can be calculated and used as the similarity.

[0090] Then, appropriate weights are assigned to the Cangjie code similarity, four-corner code similarity and stroke number similarity respectively. The weight setting can be adjusted according to actual conditions and experience, and the sum is calculated as the comprehensive similarity. If the comprehensive similarity is greater than the preset similarity threshold, such as 90%, the characters in the address information that failed to match are modified.

[0091] In one embodiment of the present application, the correcting of the pronunciation of each character of the address information includes:

[0092] Convert each character of the address information into pinyin to obtain a pinyin sequence;

[0093] Determine whether there is a preset number of repeated pinyin combinations in the pinyin sequence;

[0094] If so, the multiple Chinese characters corresponding to the repeated pinyin combination are retained as one Chinese character.

[0095] Optionally, the address information can be converted from text to pinyin, and the pinyin can be detected to determine whether there are preset repeated pinyin combinations, such as lulululu, whose repeated pinyin combination is lu, and its length is greater than 3, which indicates that the input is incorrect. At this time, the Chinese characters corresponding to the pinyin can be modified to the same character, for example, from "路路路" to "路", so as to remove repeated Chinese characters.

[0096] In an embodiment of the present application, the element to be matched includes geographical information, and matching the element to be matched with feature information of a preset point of interest includes:

[0097] Querying a preset region relationship dictionary to determine whether there is a parent-child relationship between the region information and the region information of the preset point of interest;

[0098] If a parent-child relationship exists, the match is consistent.

[0099] Optionally, the region information in the point of interest (POI) and the query is usually organized and stored in a JSON data structure. When there is region information in both the element to be matched and the preset POI, the provincial, municipal, and district Json dictionary can be queried to determine whether there is a parent-child relationship between the region information in the address information and the region information in the preset POI. If so, the match is consistent; if not, the match is inconsistent.

[0100] It should be noted that the parent-child relationship refers to the subordinate relationship in the administrative region in the scenario of region information. For example, "Shanghai" is the parent level and "Huangpu District" is the child level. If the region information of the POI and the query can correspond in this hierarchical structure and form a reasonable parent-child relationship (such as the POI contains "Shanghai" and the query contains "Huangpu District", and there is a parent-child association from the perspective of the administrative region), it meets the condition of "there is a parent-child relationship".

[0101] In an embodiment of the present application, the element to be matched includes logical sub-points. The matching of the element to be matched with the characteristic information of the preset POI includes:

[0102] Matching the logical sub-point with the logical sub-point of the preset POI;

[0103] If a continuous preset number of characters in the logical sub-point match successfully, the match is consistent.

[0104] It should be noted that a logical sub-point refers to an element unit with a specific logical meaning obtained after parsing the address information through a name structuring model. It can be understood as a segment with relatively independent logic and semantics in the address information, which may be a specific functional area, a specific building unit, a site with a specific purpose, etc. in the address. For example, in the address of a large shopping mall "XX Road, XX District, XX City, XX No., Basement 1, Food Street of XX Shopping Mall", "Food Street" can be regarded as a logical sub-point.

[0105] Optionally, the logical sub-points parsed from the address information are compared one by one with the logical sub-points included in the preset POI. For example, the logical sub-point in the address information is "Library Study Room", and there are logical sub-points such as "Library Borrowing Area" and "Library Study Room" in the preset POI. Then, the "Library Study Room" in the address will be compared with these two logical sub-points in the preset POI respectively. When a continuous preset number of characters in the logical sub-point of the address information can completely match those in the logical sub-point of the preset POI, such as 2 characters, it is considered that these two logical sub-points match consistently. It should be noted that the preset number of characters can be adjusted according to the actual situation and specific requirements.

[0106] In an embodiment of the present application, the element to be matched includes a core word, and the matching of the element to be matched with the feature information of a preset point of interest includes:

[0107] Construct an inverse document frequency dictionary, where the inverse document frequency dictionary includes inverse document frequency values corresponding to a number of points of interest;

[0108] Based on the inverse document frequency dictionary, calculate the term frequency - inverse document frequency scores corresponding to the core word and the core words in the preset points of interest respectively;

[0109] Determine the similarity between the term frequency - inverse document frequency score corresponding to the core word and the term frequency - inverse document frequency scores corresponding to the core words in the preset points of interest;

[0110] When the similarity is greater than a preset similarity threshold, the matching is consistent.

[0111] Optionally, a large amount of POI information can be collected, and this information can include text content such as the name, address, and description of the POI. Perform word segmentation on the collected POI information, and count the frequency of each word in all POI documents and the number of POI documents containing the word. According to the above statistical results, calculate the IDF value of each word according to the calculation formula of the inverse document frequency. Organize each word and its corresponding IDF value into a dictionary structure to obtain the inverse document frequency dictionary. Calculate the term frequency of the core word in the query request. And based on the inverse document frequency dictionary, determine the inverse document frequency value of the core word, and then based on the term frequency and the inverse document frequency value, obtain the tf - idf score of the core word. Similarly, for the core words in the preset points of interest, calculate the tf - idf scores according to the above method, and then combine the address information and the tf - idf scores of all the core words in the preset points of interest into a vector respectively, and calculate the similarity of the two vectors based on algorithms such as cosine similarity and Jaccard similarity. When the similarity is greater than the preset similarity threshold, the matching is consistent.

[0112] In an embodiment of the present application, the element to be matched includes a category word, and the matching of the element to be matched with the feature information of a preset point of interest includes:

[0113] Calculate the term frequency - inverse document frequency scores of the category word and the category words of the preset points of interest respectively;

[0114] Based on the term frequency - inverse document frequency scores, match the category word and the category words of the preset points of interest;

[0115] Determine the fingerprint codes corresponding to the category words with failed matching respectively;

[0116] Determine the similarity of the fingerprint codes;

[0117] If the similarity is greater than a preset similarity threshold, modify the term frequency-inverse document frequency score corresponding to the category word with a matching failure.

[0118] Optionally, first calculate the term frequency-inverse document frequency scores of the category words in the address information and the term frequency-inverse document frequency scores of the category words of the preset points of interest respectively based on the inverse document frequency dictionary. Then, combine the tf-idf scores of all the category words in the address information and the preset points of interest into a vector respectively, and calculate the similarity of the two vectors based on algorithms such as cosine similarity and Jaccard similarity. When the similarity is greater than the preset similarity threshold, the match is consistent. Otherwise, the match fails. It should be noted that due to the existence of glyph differences, for example, "Beijing City" and "Beijing" have the same semantics but different word forms, and tf-idf will regard them as different words and cannot directly match. At this time, the part with a matching failure can be calculated for similarity through the SimBERT model. SimBERT is based on a pre-trained language model and can understand the semantic information of words and sentences. That is, input the category word with a matching failure into the SimBERT model, and perform fingerprint encoding through the SimBERT model to obtain the SimHash fingerprint of each category word, and judge their similarity by comparing the Hamming distance between the fingerprints.

[0119] In the embodiments of the present application, by performing error correction processing on the address information, problems such as spelling mistakes and non-standard expressions input by users can be effectively identified and corrected, so that subsequent information processing is based on accurate data, greatly improving the usability of the address information. Using the name structuring model to perform structure parsing on the address information can accurately extract key elements such as regional words, core words, category words, and logical sub-points, comprehensively and deeply understand the connotation of the address information, and provide richer and more valuable information for subsequent matching with the feature information of the preset points of interest. Matching the elements to be matched obtained based on the in-depth parsing with the feature information of the preset points of interest can, compared with the traditional simple matching method, judge the relevance between the address information and the existing points of interest from multiple dimensions and more meticulously, thus significantly improving the matching success rate and enabling more eligible new points of interest to be discovered and supplemented. The entire process realizes the automated processing from user input to point of interest supplementation, greatly reducing manual intervention, saving time and labor costs, being able to quickly respond to a large number of user input query requests, timely supplement new points of interest to the target map, and enabling the map information to be updated and improved more quickly, meeting the user's requirements for the timeliness and richness of the map information.

[0120] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0121] In one embodiment, a map point of interest supplementing device is provided, which corresponds one-to-one with the map point of interest supplementing method in the above embodiment. As Figure 2 shown, the map point of interest supplementing device includes a query request obtaining unit 10, an error correction processing unit 20, a to-be-matched element obtaining unit 30, a matching unit 40, and a point of interest supplementing unit 50. The detailed description of each functional module is as follows:

[0122] The query request obtaining unit 10 is configured to obtain a query request input by a target user, and the query statement includes address information;

[0123] The error correction processing unit 20 is configured to perform error correction processing on the address information;

[0124] The to-be-matched element obtaining unit 30 is configured to perform structural analysis on the address information after error correction processing to obtain to-be-matched elements;

[0125] The matching unit 40 is configured to match the to-be-matched elements with the feature information of preset points of interest;

[0126] The point of interest supplementing unit 50 is configured to, if the matching is successful, use the address information as a new point of interest and supplement it to the target map.

[0127] In one embodiment of the present application, the error correction processing unit 20 is further configured to:

[0128] Perform error correction on the glyphs of each character of the address information;

[0129] Perform error correction on the pronunciation of each character of the address information.

[0130] In one embodiment of the present application, the error correction processing unit 20 is further configured to:

[0131] Match the address information with the preset points of interest character by character;

[0132] If the character in the address information fails to match the character in the preset point of interest, determine whether the characters before and after the character that fails to match are successfully matched;

[0133] If the matching is successful, determine the similarity between the address information and the character that fails to match in the preset point of interest;

[0134] If the similarity is greater than a preset similarity threshold, modify the character that fails to match.

[0135] In one embodiment of the present application, the error correction processing unit 20 is further configured to:

[0136] Determine the Cangjie code similarity between the address information and the characters that failed to match in the preset point of interest;

[0137] Determine the four-corner code similarity between the address information and the characters that failed to match in the preset point of interest;

[0138] Determine the stroke count similarity between the address information and the characters that failed to match in the preset point of interest;

[0139] Based on the Cangjie code similarity, four-corner code similarity, and stroke count similarity, determine the comprehensive similarity.

[0140] In an embodiment of the present application, the error correction processing unit 20 is further configured to:

[0141] Convert each character of the address information into pinyin to obtain a pinyin sequence;

[0142] Determine whether there are preset repeated pinyin combinations in the pinyin sequence;

[0143] If so, retain multiple Chinese characters corresponding to the repeated pinyin combination as one Chinese character.

[0144] In an embodiment of the present application, the error correction processing unit 20 is further configured to:

[0145] Query a preset regional relationship dictionary to determine whether there is a parent-child relationship between the regional information and the regional information of the preset point of interest;

[0146] If there is a parent-child relationship, the match is consistent.

[0147] In an embodiment of the present application, the error correction processing unit 20 is further configured to:

[0148] Match the logical sub-point with the logical sub-point of the preset point of interest;

[0149] If a continuous preset number of characters in the logical sub-point match successfully, the match is consistent.

[0150] In an embodiment of the present application, the error correction processing unit 20 is further configured to:

[0151] Construct an inverse document frequency dictionary, where the inverse document frequency dictionary includes inverse document frequency values corresponding to several points of interest;

[0152] Based on the inverse document frequency dictionary, calculate the term frequency-inverse document frequency scores corresponding to the core word and the core word in the preset point of interest respectively;

[0153] Determine the similarity between the term frequency-inverse document frequency score corresponding to the core word and the term frequency-inverse document frequency score corresponding to the core word in the preset point of interest;

[0154] When the similarity is greater than a preset similarity threshold, the match is consistent.

[0155] In an embodiment of the present application, the error correction processing unit 20 is further configured to:

[0156] Calculate the term frequency-inverse document frequency scores of the category word and the category word of the preset point of interest respectively;

[0157] Based on the term frequency-inverse document frequency scores, match the category word and the category word of the preset point of interest;

[0158] Determine the fingerprint codes corresponding to the category words with failed matches respectively;

[0159] Determine the similarity of the fingerprint codes;

[0160] If the similarity is greater than a preset similarity threshold, modify the term frequency-inverse document frequency scores corresponding to the category words with failed matches.

[0161] In the embodiment of the present application, by performing error correction processing on the address information, problems such as spelling mistakes and non-standard expressions input by the user can be effectively identified and corrected, enabling subsequent information processing to be based on accurate data, and greatly improving the usability of the address information. Using the name structuring model to perform structural analysis on the address information can accurately extract key elements such as geographical words, core words, category words, and logical sub-points, comprehensively and deeply understand the connotation of the address information, and provide richer and more valuable information for subsequent matching with the feature information of the preset point of interest. Matching the elements to be matched obtained through in-depth analysis with the feature information of the preset point of interest can, compared with the traditional simple matching method, judge the relevance between the address information and the existing points of interest from multiple dimensions and more meticulously, thereby significantly improving the matching success rate and enabling more eligible new points of interest to be discovered and supplemented. The entire process realizes automated processing from user input to point of interest supplementation, greatly reducing manual intervention, saving time and labor costs, being able to quickly respond to a large number of query requests input by users, promptly supplementing new points of interest to the target map, and enabling the map information to be updated and improved more quickly to meet the user's requirements for the timeliness and richness of the map information.

[0162] For the specific limitations of the map point of interest supplementation device, reference can be made to the limitations on the map point of interest supplementation method in the above text, which will not be elaborated here. Each module in the above map point of interest supplementation device can be implemented in whole or in part through software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0163] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as shown in Figure 3 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a method for supplementing map points of interest is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0164] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the method for supplementing map points of interest as described above are implemented.

[0165] In an embodiment of the application, a readable storage medium is provided. The readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the method for supplementing map points of interest as described above are implemented.

[0166] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0167] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0168] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for supplementing points of interest on a map, characterized in that: The method comprises: Acquire a query request input by a target user, wherein the query request includes address information; performing error correction processing on the address information; Performing structural analysis on the address information after error correction to obtain elements to be matched, wherein the elements to be matched include at least one of a regional word, a core word, a category word, and a logical sub-point; Matching the element to be matched with feature information of a preset point of interest; If the match is successful, the address information is used as a new point of interest and added to the target map.

2. The method for supplementing points of interest on a map according to claim 1, characterized in that: The error correction process of the address information includes: Correcting the glyphs of each character of the address information; Correct the pronunciation of each character in the address information.

3. The method for supplementing points of interest on a map according to claim 1, characterized in that: The step of correcting the glyphs of the characters in the address information comprises: Matching the address information with the preset points of interest character by character; If the characters in the address information fail to match the characters in the preset point of interest, determining whether the characters before and after the characters that failed to match are matched successfully; If the match is successful, determining the similarity between the address information and the characters in the preset point of interest where the match failed; If the similarity is greater than a preset similarity threshold, the matching failure character is modified.

4. The method for supplementing points of interest on a map according to claim 3, characterized in that: The determining the similarity between the address information and the characters in the preset point of interest where the matching fails, includes: Determining the Cangjie code similarity between the address information and the characters in the preset point of interest where the matching fails; Determine the four-corner code similarity between the address information and the matching failure character in the preset point of interest; Determine the similarity in number of strokes between the address information and the character whose matching fails in the preset point of interest; Based on the Cangjie code similarity, the four-corner code similarity and the stroke number similarity, a comprehensive similarity is determined.

5. The method for supplementing map points of interest as claimed in claim 2, characterized in that: The correcting of the pronunciation of each character of the address information comprises: Convert each character of the address information into pinyin to obtain a pinyin sequence; Determine whether there is a preset number of repeated pinyin combinations in the pinyin sequence; If so, the multiple Chinese characters corresponding to the repeated pinyin combination are retained as one Chinese character.

6. The method for supplementing points of interest on a map according to claim 1, characterized in that: The element to be matched includes geographical information, and matching the element to be matched with feature information of a preset point of interest includes: Querying a preset region relationship dictionary to determine whether there is a parent-child relationship between the region information and the region information of the preset point of interest; If a parent-child relationship exists, the match is consistent.

7. The method for supplementing points of interest on a map according to claim 1, characterized in that: The element to be matched includes a logical sub-point, and matching the element to be matched with feature information of a preset point of interest includes: Matching the logical sub-point with the logical sub-point of the preset point of interest; If a preset number of characters in the logical subpoint are matched successfully, the match is consistent.

8. The method for supplementing points of interest on a map according to claim 1, characterized in that: The element to be matched includes a core word, and matching the element to be matched with feature information of a preset point of interest includes: Constructing an inverse document frequency dictionary, wherein the inverse document frequency dictionary includes inverse document frequency values ​​corresponding to a plurality of interest points; Based on the inverse document frequency dictionary, respectively calculating the word frequency-inverse document frequency scores corresponding to the core words and the core words in the preset interest points; Determine the similarity between the word frequency-inverse document frequency score corresponding to the core word and the word frequency-inverse document frequency score corresponding to the core word in the preset interest point; When the similarity is greater than a preset similarity threshold, the match is consistent.

9. The method for supplementing points of interest on a map according to claim 1, characterized in that: The element to be matched includes a category word, and matching the element to be matched with feature information of a preset point of interest includes: Calculating the word frequency-inverse document frequency scores of the classifier and the classifier of the preset interest point respectively; Based on the word frequency-inverse document frequency score, matching the classifier with the classifier of the preset point of interest; Determine the fingerprint codes corresponding to the category words that failed to match respectively; Determining the similarity of the fingerprint codes; If the similarity is greater than a preset similarity threshold, the word frequency-inverse document frequency score corresponding to the category word that failed to match is modified.

10. A device for supplementing points of interest on a map, characterized in that: The device comprises: A query request acquisition unit, used to acquire a query request input by a target user, wherein the query statement includes address information; An error correction processing unit, used for performing error correction processing on the address information; A to-be-matched element acquisition unit is used to perform structural analysis on the address information after error correction processing to obtain the to-be-matched element; A matching unit, used to match the element to be matched with feature information of a preset point of interest; The point of interest supplement unit is used to take the address information as a new point of interest and supplement it into the target map if the match is successful.