A method, device and equipment for identifying geographic location

By calculating the similarity parameters of the description information of the geographical location identification, the problem of difficult to identify different identifiers of the same geographical location in the prior art is solved, and a more accurate geographical location identification matching and user service experience are achieved.

CN112966125BActive Publication Date: 2025-05-23CHINA CONSTRUCTION BANK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110472704.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-29
Publication Date
2025-05-23
Estimated Expiration
2041-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify different geographical location identifiers for the same geographical location, resulting in the inability to provide effective services.

Method used

By obtaining text containing different geographical location identifiers, finding and extracting description information describing these identifiers, and calculating similarity parameters between the description information. If the similarity parameters are greater than the set threshold, it is determined that these identifiers are identified as identifiers describing the same geographical location.

Benefits of technology

There is no need to pre-construct a data table of geographical location names. By directly comparing the information in the text, the accuracy of similarity comparison between geographical location identifiers is improved and the user experience is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112966125B_ABST
    Figure CN112966125B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a geographic location identification method, device and equipment, which are applied to the field of big data technology. The method includes: obtaining a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; searching for first description information in the first matching text, and searching for second description information in the second matching text; calculating the similarity parameter between the first description information and the second description information; when the similarity parameter is greater than the similarity threshold, determining that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location. The above method improves the accuracy of the comparison results, facilitates the services implemented by subsequent users based on different geographic location identifiers, and optimizes the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of big data technology, and more particularly to a method, device and equipment for identifying geographic location. Background Art

[0002] With the development of science and technology, people can more easily navigate to a certain place or query information about a certain place. The premise for achieving the above effect is that the user can accurately input the name corresponding to the target geographical location, or the server can accurately identify the geographical location corresponding to the name after receiving the name input by the user. However, in news reports, user chat records and other types of notification texts, more and more geographical location information appears. This geographical location information may be different expressions for the same geographical location. As a result, when different geographical location names are received, it may not be possible to effectively determine the correct geographical location corresponding to these geographical location names, and it is therefore impossible to provide subsequent services.

[0003] At present, when solving this problem, all corresponding geographic location names are often listed in advance for each geographic location based on the experience of the management personnel. When receiving the geographic location identifier input by the user, the geographic location identifier is matched with the recorded geographic location name to obtain the desired geographic location. However, this method has high requirements for the integrity of the recorded geographic location name. If the corresponding geographic location name is not recorded in advance, the matching search cannot be completed, thereby affecting the user's subsequent use experience. Therefore, there is an urgent need for a method that can accurately identify the same identifier for the same geographic location. Summary of the invention

[0004] The purpose of the embodiments of this specification is to provide a method, device and equipment for identifying a geographic location to solve the problem of how to accurately identify the same identifier for the same geographic location.

[0005] To solve the above technical problems, an embodiment of the present specification provides a geographic location identification method, including: obtaining a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; searching for first description information in the first matching text, and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; calculating a similarity parameter between the first description information and the second description information; the size of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; when the similarity parameter is greater than a similarity threshold, determining that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location.

[0006] The embodiment of the present specification also proposes a geographic location identification device, including: a matching text acquisition module, used to obtain a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; a description information search module, used to search for first description information in the first matching text, and search for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; a similarity parameter calculation module, used to calculate a similarity parameter between the first description information and the second description information; the size of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; a determination module, used to determine that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location when the similarity parameter is greater than a similarity threshold.

[0007] The embodiment of the present specification also proposes a geographic location identification device, including a memory and a processor; the memory is used to store computer program instructions; the processor is used to execute the computer program instructions to implement the following steps: obtaining a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; searching for first description information in the first matching text, and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; calculating a similarity parameter between the first description information and the second description information; the size of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; when the similarity parameter is greater than a similarity threshold, determining that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location.

[0008] As can be seen from the technical solutions provided in the embodiments of this specification above, after obtaining texts respectively containing different geographical location identifiers, the embodiments of this specification search for description information describing the geographical location identifiers from these texts, and calculate the similarity parameters between these description information based on the categories corresponding to the description information, quantitatively describe the similarity between geographical location identifiers, and finally can determine whether different geographical location identifiers are used to describe the same geographical location based on the comparison result between the similarity parameter and the similarity threshold. The above method does not need to pre-construct a data table containing different geographical location names, and directly compares through the information contained in the text to determine the similarity degree between geographical location identifiers, improves the accuracy of the comparison result, facilitates the services implemented by users based on different geographical location identifiers in the subsequent process, and optimizes the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 It is a flowchart of a geographical location recognition method according to an embodiment of this specification;

[0011] Figure 2 It is a module diagram of a geographical location recognition device according to an embodiment of this specification;

[0012] Figure 3 It is a structural diagram of a geographical location recognition device according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0014] To solve the above technical problems, a geographical location recognition method according to an embodiment of this specification is introduced. The execution subject of the geographical location recognition method is a geographical location recognition device, and the geographical location recognition device includes but is not limited to a server, an industrial control computer, a PC, etc. As Figure 1 shown, the geographical location recognition method may include the following specific implementation steps.

[0015] S110: Obtain a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier.

[0016] The first matching text and the second matching text may be different text information, such as news reports, community notifications, online posts, or chat records. Due to differences in the format, audience, and writing requirements of different matching texts, different identifiers may be used to describe the same geographical location in these matching texts. For example, for the same geographical location, there may be a formally registered name, a name based on the geographical location, and a name used for promotional descriptions. In formal reports such as news, formally registered names may be used for descriptions, while general online posts or chat records may use more common names for promotional descriptions. When different descriptions of the same geographical location are used in different texts, it is easy to mistakenly believe that the two texts describe different geographical locations, thereby causing interference in subsequent applications. Therefore, it is necessary to determine whether different geographical location identifiers are the same geographical location.

[0017] In some practical applications, the same geographic location may be represented by different description methods in the same text. In this case, the first matching text and the second matching text may be the same text, but at least two different geographic location identifiers may be determined in the text as the first geographic location identifier and the second geographic location identifier, respectively.

[0018] Generally, the first matching text and the second matching text contain more text information, and these text information contain text information describing the corresponding geographic location identifiers. In the case where it is impossible to directly determine whether the geographic locations are the same through the geographic location identifiers, it is possible to indirectly determine whether the geographic location identifiers correspond to the same geographic location based on the similarity of the content corresponding to these description information.

[0019] The first matching text contains a first geographic location identifier, and the second matching text contains a second geographic location identifier. The first geographic location identifier and the second geographic location identifier may be descriptions of the geographic location in the form of a name, a number, etc. For example, when the geographic location is the location of a real estate project, the first geographic location identifier and / or the second geographic location identifier include an identifier corresponding to the real estate project. Specifically, assuming that they correspond to the same real estate project, the first geographic location identifier may be "Future Happy Garden" and the second geographic location identifier may be "MAX Future". However, it may not be possible to determine whether the first geographic location identifier and the second geographic location identifier are identifiers corresponding to the same geographic location directly based on the first geographic location identifier and the second geographic location identifier, so it is necessary to determine them based on certain steps.

[0020] In some embodiments, the method for obtaining the first geographic location identifier may be to first perform segmentation on the first matching text to obtain at least one first text segmentation. The specific segmentation method can be set based on the needs of the actual application and will not be described in detail here. After obtaining these first text segmentations, it is possible to determine in turn whether these first text segmentations are words used to describe the geographic location. The specific identification method may pre-train the corresponding model to identify the words used to describe the geographic location in the segmentation, or may pre-set some specific words. When these words are included in the segmentation, it can be determined that these segmentations are geographic location identifiers. In actual applications, geographic location identifiers may also be identified in other ways, which are not limited to this. Correspondingly, the method for obtaining the second geographic location identifier may be to perform segmentation on the second matching text to obtain at least one second text segmentation, and then identify the second geographic location identifier from the second text segmentation.

[0021] Since the first geographic location identifier and the second geographic location identifier are the same, it can be obviously considered that the first geographic location identifier and the second geographic location identifier refer to the same geographic location. Therefore, preferably, based on the geographic location identification method implemented in the embodiments of this specification, the first geographic location identifier and the second geographic location identifier are different identifiers.

[0022] S120: Searching for first description information in the first matching text, and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier.

[0023] The first description information is information describing the first geographic location identifier; the second description information is information describing the second geographic location identifier. The above description information can be a detailed explanation of the geographic location identifier, or text information that is strongly associated with the geographic location identifier.

[0024] In some embodiments, the first description information and the second description information may be at least one of regional information, longitude and latitude information, building name information, road name information, road number information, building type information, property information, building age information, resident information, and environmental information. The regional information may be the administrative region where the geographical location is located. For example, it may be Haidian District, Beijing. In practical applications, the regional information may also be further refined and classified. The longitude and latitude information may be the longitude and latitude corresponding to the geographical location. The building name information may be the name of a specific real estate project. The road name information may be the name of the road where the geographical location is located, and the road number information may be the specific number of the geographical location on the road where it is located. Specifically, it may be No. 158, Nanjing Road. The building type information may be the specific type of the building at the geographical location, such as ordinary residence, apartment, villa, dormitory, etc. The property information may be the information about the management of the location of the geographical location, such as the property name, the property management duration, etc. The building age information may be the time when the building at the geographical location was built. The resident information may be the information such as the number of households and the number of residents in the building at the geographical location. The environmental information may be the information about the environment corresponding to the geographical location, such as the greening rate, the altitude, etc.

[0025] The specific process of identifying the first description information and the second description information may be determined based on the relevance between sentences in the text. For example, after finding the corresponding geographical location identifier in the matching text, the text may be semantically analyzed in combination with the identified geographical location identifier to determine the text that has a strong association with the geographical location identifier as the corresponding description information. For example, when it is detected that the first geographical location identifier or the second geographical location identifier is juxtaposed with a conjunction such as "is" or "of", the text after the conjunction may be used as the first description information or the second description information.

[0026] In other embodiments, since the first description information and the second description information are used to describe the geographical location, information with a strong association with the geographical location may also be directly searched for in the first matching text or the second matching text as the first description information and the second description information. For example, some geographical location formats may be preset, such as text containing specific type of vocabulary such as streets, as the first description information or the second description information. The specific identification method may be set according to the requirements of practical applications and will not be elaborated here.

[0027] In some embodiments, the first description information and the second description information may also be from the first geographical location identifier and the second geographical location identifier themselves. Since the first geographical location identifier and the second geographical location identifier may have some key participles for describing the geographical location when describing the geographical location, such key participles can be directly extracted from the geographical location identifier to identify the corresponding geographical location.

[0028] Correspondingly, the description information can be used to describe the word occurrence frequency in the geographical location identifier, so as to determine the keywords in the geographical location identifier based on these frequencies, and then make a more accurate definition based on the keywords.

[0029] Specifically, the first geographical location identifier can be segmented to obtain at least one first location participle, and then the first word frequency of the first location participle in the first matching text and the first inverse word frequency in the sample text can be determined respectively. Correspondingly, for the second geographical location identifier, the second geographical location identifier can also be segmented to obtain at least one second location participle, and then the second word frequency of the second location participle in the second matching text and the second inverse word frequency in the sample text can be determined respectively.

[0030] The first word frequency can be used to represent the total frequency of the first location participle in all the words of the first matching text, and the first inverse word frequency can be used to represent the frequency of the first location participle in other sample texts. Correspondingly, the second word frequency can be used to represent the total frequency of the second location participle in all the words of the second matching text, and the second inverse word frequency can be used to represent the frequency of the second location participle in other sample texts.

[0031] When the frequency of a certain word in the current text is higher and the frequency in other texts is lower, it means that the word has a higher representativeness for this text, that is, the word is more likely to be a keyword and has the effect of prominently describing the geographical location. Correspondingly, when comparing different geographical location identifiers based on the keywords, the accuracy of the comparison results obtained is higher. And when the frequency of a word is higher in both the current text and other texts, the word is more likely to be a relatively common descriptive word, such as words like "of", "and", etc., which usually do not have the effect of representing the geographical location.

[0032] In a specific example, the calculation of the word frequency and the inverse word frequency can be implemented based on the TF-IDF method. Specifically, the formula can be used to calculate the first word frequency, where tf i,1 is the first word frequency, and n i,1is the number of occurrences of the first position participle in the first matching text, ∑ k n k,1 is the number of all words in the first matching text; using the formula Calculate the first inverse word frequency, where idf 1 is the first inverse word frequency, |D| is the number of all geographic location identifiers in the sample text, |{j:t 1 ∈d 1}| is the number of geographic location identifiers containing the first position participle in the sample text. Correspondingly, the formula can also be used Calculate the second word frequency, where tf i,2 is the second word frequency, n i,2 The number of occurrences of the second position participle in the second matching text, ∑ k n k,2 is the number of all words in the second matching text; using the formula idf 2 = Calculate the second inverse word frequency, where idf 2 is the second inverse word frequency, |D| is the number of all geographic location identifiers in the sample text, |{j:t 2 ∈d 2}| is the number of geographic location identifiers containing the second position participle in the sample text.

[0033] By determining the location keywords and location non-keywords, it is possible to determine what is emphasized in the geographic location identifier, so that the geographic location identifier itself can be more accurately compared with other geographic location identifiers, thereby improving the accuracy of geographic location judgment.

[0034] The above examples are only schematic illustrations of the method of obtaining the first description information or the second description information. In actual applications, the first description information and the second description information may be defined in other ways, and the first description information and the second description information may be obtained in other ways accordingly. The examples are not limited to the above examples and will not be repeated here.

[0035] S130: Calculate a similarity parameter between the first description information and the second description information; the magnitude of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier.

[0036] After obtaining the first description information and the second description information, the similarity between the description information can be quantitatively determined by calculating the similarity parameter between the description information, so that it can be determined whether different geographic location identifiers are identifiers corresponding to the same geographic location based on the similarity parameter.

[0037] In some embodiments, after obtaining the first description information and the second description information, the information categories to which they correspond may be determined respectively. The information categories may be used to represent different classifications to which the information corresponds. Since the description information may describe the geographic location from different angles, using the first description information and the second description information from different angles to determine the similarity of the geographic location identifiers may affect the accuracy of the determination result. Therefore, before determining the degree of similarity, the category to which the description information corresponds may be determined first, so that the description information of the same information category may be compared to determine whether it is the same geographic location, thereby improving the accuracy of the determination result.

[0038] Correspondingly, the information category may also correspond to the categories to which different types of information belong, such as regional information, longitude and latitude information, building name information, road name information, road number information, building type information, property information, building age information, resident information, and environmental information.

[0039] A specific method for determining the information category may be to train a corresponding classifier model based on pre-labeled sample data, thereby using the trained classifier model to classify descriptive information of different information categories. The specific training method and classification process may be set based on actual application conditions and will not be described in detail here.

[0040] In some embodiments, when the first word frequency, the first reverse word frequency, the second word frequency, and the second reverse word frequency are determined, when determining the information category, the first location keyword and the first location non-keyword in the first geographic location identifier may be determined based on the first word frequency and the first reverse word frequency, and the second location keyword and the second location non-keyword in the second geographic location identifier may be determined based on the second word frequency and the second reverse word frequency. Location keywords may be segmented words with strong geographic location relevance, and may be representative descriptions of information corresponding to the geographic location. Location non-keywords may be words with unclear meanings, such as conjunctions, common nouns, etc. By distinguishing between keywords and non-keywords, comparisons can be more effectively achieved in subsequent processes, and geographic location identifiers corresponding to the same geographic location can be better determined.

[0041] By determining the information categories corresponding to different descriptive information, the descriptive information under the same information category can be used for comparison in the subsequent judgment process, thereby further improving the accuracy of the information comparison process and achieving effective identification of the identifiers of the same geographical location.

[0042] After determining the categories of each description information, a corresponding similarity parameter can be calculated according to the description information for different information categories. The size of the similarity parameter is used to indicate the similarity between the first description information and the second description information. Different rules can be set for calculating the corresponding similarity for description information of different categories.

[0043] Specifically, the first information field and the second information field corresponding to each same information category may be divided from the first description information and the second description information, and the category similarity parameters corresponding to each information category may be calculated based on the first information field and the second information field, and then the similarity parameters between the first description information and the second description information may be obtained by combining the category similarity parameters.

[0044] Table 1 below shows an example of a calculation method. There are corresponding comparison methods and judgment rules for different fields. For example, for regional information, the similarity parameter can be determined based on the distance between regions; for building names, the similarity can be calculated by judging whether the building names are the same.

[0045]

[0046] Table 1

[0047] It should be noted that the latitude and longitude information category in the above table has no weight value assigned. When the longitude and latitude corresponding to different geographic location identifiers are the same, it can be determined without a doubt that the two geographic location identifiers correspond to the same geographic location. When the longitude and latitude corresponding to different geographic location identifiers are different, the two geographic location identifiers generally correspond to different geographic locations. That is, if the first matching text and the second matching text both contain longitude and latitude information, it can be directly determined based on the longitude and latitude information whether the corresponding positions of the geographic location identifiers are the same, without the need to assign weights to them.

[0048] Of course, the above table is only an exemplary introduction to some judgment rules. In actual applications, other judgment rules can be set based on different information categories. It is not limited to the above examples and will not be repeated here.

[0049] In some embodiments, corresponding weight values ​​may be set for different information categories according to their importance. When calculating the similarity parameters, the similarity parameters between the first description information and the second description information may be obtained based on the weight values ​​corresponding to each information category and the category similarity parameters. For example, in general, the possibility of duplication of road names, roads, and area names is small, and these information categories may be assigned higher weight values; and for property information, building age information, etc., there is a high probability that buildings in different geographical locations have the same property or were built in the same year, which may easily interfere with the correct judgment process, and a smaller weight value may be assigned. The above examples are only schematic introductions to the process of weight value assignment. In actual applications, corresponding weight values ​​may be assigned for different situations, not limited to the above examples.

[0050] Table 1 above also exemplarily gives the weight values ​​set for different information categories, which is only a reference for the allocation of weight values. The weight values ​​can also be adjusted in actual applications.

[0051] In some implementations, if corresponding keywords are identified for the first geographic location identifier and the second geographic location identifier, different calculations can be performed based on the relationship between the keywords. Specifically, when the first location keyword and the second location keyword are the same, the formula Y=X can be used. 1 +X 2 +X 6 +X 8 +X 9 +X 10 +X 11 -X 4 -X 5 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 2 is the similarity parameter corresponding to the keyword, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information, X 4 is the frequency parameter corresponding to the keyword, X 5is the word frequency parameter corresponding to the non-keyword. When the first position keyword and the second position keyword are the same, the influence of the word frequency of the position keyword in different texts can be referred to, so that the word frequency parameter of the keyword and the word frequency parameter of the non-keyword can be introduced into the formula for reference to the calculation process.

[0052] In the case where the first position keyword and the second position keyword are different, the formula Y=X is used. 1 +X 6 +X 8 +X 9 +X 10 +X 11 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information.

[0053] When the keywords are the same or different, different methods are used to calculate the similarity parameters, taking into account the impact of different situations and ensuring the accuracy of the judgment results.

[0054] In other embodiments, if the keywords and non-keywords in the geographic location identifier are not determined in advance, the final similarity parameter can also be directly calculated by the similarity between the first description information and the second description information under different categories. For example, after determining the similarity parameters between the description information of each information category in turn, based on the weight values ​​corresponding to the different information categories, these similarity parameters can be directly accumulated as the total similarity parameter. The obtained total similarity parameter can be applied to the calculation in subsequent steps.

[0055] S140: If the similarity parameter is greater than a similarity threshold, determine that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location.

[0056] After the similarity parameter is calculated, the similarity parameter can be compared with a similarity threshold. The similarity threshold can be a pre-set parameter value used to represent the minimum value of the similarity parameter between geographic location identifiers corresponding to the same geographic location.

[0057] The similarity threshold can be set based on human experience by roughly judging the proximity between the description information corresponding to two geographical location identifiers to determine the similarity threshold. It can also be to obtain a certain amount of sample data, which can be location identifiers and the description information corresponding to the location identifiers, and add corresponding labels to these sample data to indicate whether the geographical locations corresponding to these sample data are the same. Then, the method in the embodiments of this specification can be used to calculate the similarity parameter for each sample data. Finally, the minimum value of the similarity parameter between the geographical location identifiers satisfying the same geographical location is determined in combination with the similarity parameter and the label corresponding to the sample data, that is, the similarity threshold. In practical applications, the similarity threshold can also be determined by other means, which will not be elaborated here.

[0058] Therefore, when the similarity parameter is greater than the similarity threshold, it can be determined that the first geographical location identifier and the second geographical location identifier are identifiers describing the same geographical location.

[0059] Correspondingly, when the similarity parameter is not greater than the similarity threshold, it indicates that there is no strong correlation between the first geographical location identifier and the second geographical location identifier, and it can be determined that the first geographical location identifier and the second geographical location identifier are identifiers describing different geographical locations. This property can also be used to distinguish other geographical location identifiers in subsequent processes.

[0060] In some embodiments, after determining that the first geographical location identifier and the second geographical location identifier are identifiers describing the same geographical location, the first geographical location identifier and the second geographical location identifier can be registered in the same data table. In subsequent application processes, by checking whether another geographical location identifier is included in the data table corresponding to a certain geographical location identifier, it can be determined whether these two geographical location identifiers are identifiers corresponding to the same geographical location, which is convenient for subsequent applications.

[0061] If it is determined that the first geographical location identifier and the second geographical location identifier are different geographical locations, it can first be determined whether one of the geographical location identifiers has been registered in a certain data table. If it has been registered, a new data table can be created to record the other geographical location identifier, and these two geographical location identifiers can be distinguished by different data tables; if neither has been registered, new data tables can be created respectively to record these two geographical location identifiers.

[0062] When performing the above operations, as the number of created data tables increases, operations such as merging and splitting these data tables can also be performed in subsequent processes, so as to associate the data tables based on the relationship between geographical location identifiers, avoiding the process of repeatedly creating redundant data tables, and enabling effective utilization of the discrimination results.

[0063] It should be noted that the method in the above embodiment is only an exemplary method for determining the geographic location identifiers in two matching texts. In actual applications, based on the implementation process of the above method, the number of matching texts can be expanded to more than two to achieve the determination of identifiers corresponding to the same geographic location in multiple matching texts. It is also possible to compare each two geographic location identifiers in turn, and based on the correlation between the comparison results, combine all the comparison results to obtain the recognition results between multiple geographic location identifiers. The specific implementation process can be directly implemented based on the above method steps, and will not be repeated here.

[0064] The execution process of the above method is explained using a specific scenario example. In this scenario example, it is assumed that there are two reports on the same real estate, which are marked as report A and report B. In report A, the name of the real estate is registered as "No. 1 Commercial and Office Building, Juhan Plaza", and the opening time, building area, greening rate, and the area where the real estate is located are introduced; in report B, the name of the real estate is registered as "Excellent Zhonghuan", and the building area, corresponding road name, greening rate and property type of the real estate are introduced. After obtaining report A and report B, the execution device can first segment report A and report B respectively, and identify the information that meets the geographic location identifier, thereby obtaining "No. 1 Commercial and Office Building, Juhan Plaza" and "Excellent Zhonghuan" respectively. Further word segmentation is performed on the "Juhan Plaza No. 1 Commercial and Office Building", and the word segmentations corresponding to "Juhan Plaza No. 1 Commercial and Office Building" are "Juhan Plaza", "No. 1", "Commercial", and "Office Building". Through analysis, it can be determined that the first keyword is "Juhan Plaza" and "No. 1"; and the word segmentations corresponding to "Excellent Zhonghuan Commercial Center" are "Excellent Zhonghuan" and "Commercial Center". Through analysis, it can be determined that the second keyword is "Excellent Zhonghuan".

[0065] Afterwards, the first descriptive information for "No. 1 Commercial and Office Building, Juhan Plaza" is searched in report A, and the information categories corresponding to each first descriptive information are determined by performing semantic recognition on the first descriptive information. Specifically, the information category of the information "opened on March 16" can be classified as the opening time, and the information categories corresponding to other first descriptive information are determined in turn as the building area, greening rate and location; the second descriptive information for "Excellent Zhonghuan" is searched in report B, and the information categories corresponding to each second descriptive information are determined by performing semantic recognition on the second descriptive information. Specifically, the information category of the information "greening rate as high as 60%" can be classified as greening rate, and the information categories corresponding to other second descriptive information are determined in turn as the building area, corresponding road name and property type.

[0066] In the case of determining the information categories corresponding to different descriptive information, the similarity between different descriptive information is analyzed based on the descriptive information under the corresponding information category. In the case where the keywords of the above two geographic location identifiers are different, there is no need to consider the impact of the keyword frequency on the similarity of the two location identifiers. The similarity parameters between the descriptive information under each information category are directly calculated, and then the overall similarity parameters are accumulated. The specific calculation data is no longer introduced here. After accumulating the overall similarity parameter, by comparing the similarity parameter with the similarity threshold, it is determined that the similarity parameter is greater than the similarity threshold, indicating that the proximity between the two geographic location identifiers meets the conditions of the same geographic location, that is, the "Juhan Plaza No. 1 Commercial and Office Building" and "Excellent Zhonghuan" are the same geographic location. The two geographic location identifiers can be recorded by registering them in the same data table. If the two geographic location identifiers are found in the text in the subsequent process, the geographic location identifiers can be effectively converted.

[0067] Through the introduction of the above embodiments and scenario examples, it can be seen that after obtaining texts containing different geographic location identifiers, the method searches for description information describing the geographic location identifiers from these texts, and calculates the similarity parameters between these description information based on the categories corresponding to the description information, quantitatively describes the similarity between the geographic location identifiers, and finally determines whether different geographic location identifiers are used to describe the same geographic location based on the comparison result between the similarity parameter and the similarity threshold. The above method does not need to pre-build a data table containing different geographic location names, and determines the similarity between geographic location identifiers by directly comparing the information contained in the text, thereby improving the accuracy of the comparison results, facilitating the services implemented by users based on different geographic location identifiers in the subsequent process, and optimizing the user experience.

[0068] based on Figure 1 The corresponding geographic location identification method introduces a geographic location identification device in the embodiment of this specification. The geographic location identification device is set in the geographic location identification device. Figure 2 As shown, the geographic location identification device includes the following modules.

[0069] The matching text acquisition module 210 is used to acquire a first matching text and a second matching text; the first matching text includes a first geographic location identifier; the second matching text includes a second geographic location identifier.

[0070] The description information search module 220 is used to search for first description information in the first matching text and search for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier.

[0071] The similarity parameter calculation module 230 is used to calculate the similarity parameter between the first description information and the second description information; the size of the similarity parameter is used to represent the similarity between the first geographic location identifier and the second geographic location identifier.

[0072] The determination module 240 is configured to determine, when the similarity parameter is greater than a similarity threshold, that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location.

[0073] based on Figure 1 Corresponding to the geographic location identification method, the embodiment of this specification provides a geographic location identification device. Figure 3 As shown, the geographic location identification device may include a memory and a processor.

[0074] In this embodiment, the memory may be implemented in any appropriate manner. For example, the memory may be a read-only memory, a mechanical hard disk, a solid-state hard disk, or a USB flash drive, etc. The memory may be used to store computer program instructions.

[0075] In this embodiment, the processor may be implemented in any appropriate manner. For example, the processor may be in the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. The processor may execute the computer program instructions to implement the following steps: obtaining a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; searching for first description information in the first matching text and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; calculating a similarity parameter between the first description information and the second description information; the size of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; when the similarity parameter is greater than a similarity threshold, determining that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location.

[0076] It should be noted that the geographical location recognition method, device and equipment can be applied to the field of big data technology, and can also be applied to other technical fields, which is not restricted herein.

[0077] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0078] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0079] It can be known from the above description of the implementation mode that the technicians in this field can clearly understand that the present specification can be implemented by means of software plus the necessary first hardware platform. Based on such an understanding, the technical solution of the present specification can be essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present specification or some parts of the embodiments.

[0080] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0081] This specification can be used in many first or special computer system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0082] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0083] Although the present specification is described through embodiments, those skilled in the art will appreciate that there are many modifications and changes to the present specification without departing from the spirit of the present specification, and it is intended that the appended claims include these modifications and changes without departing from the spirit of the present specification.

Claims

1. A method for identifying a geographic location, It is characterized in that include: Get the first matching text and the second matching text; The first matching text contains a first geographical location identifier; The second matching text contains a second geographical location identifier; Searching for first description information in the first matching text, and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; Calculating a similarity parameter between the first description information and the second description information; The magnitude of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; When the similarity parameter is greater than a similarity threshold, determining that the first geographical location identifier and the second geographical location identifier are identifiers describing the same geographical location; Searching for first description information in the first matching text includes: segmenting the first geographical location identifier to obtain at least one first position segmentation; determining a first word frequency of the first position segmentation in the first matching text and a first reverse word frequency in the sample text respectively; Searching for second description information in the second matching text includes: segmenting the second geographical location identifier to obtain at least one second position segmentation; determining a second word frequency of the second position segmentation in the second matching text and a second reverse word frequency in the sample text respectively; Determining information categories corresponding to the first description information and the second description information includes: determining a first location keyword and a first location non-keyword in the first geographic location identifier based on the first word frequency and the first reverse word frequency; determining a second location keyword and a second location non-keyword in the second geographic location identifier based on the second word frequency and the second reverse word frequency; Calculating the similarity parameter between the first description information and the second description information based on the information category includes: when the first position keyword and the second position keyword are the same, using the formula Y=X 1 +X 2 +X 6 +X 8 +X 9 +X 10 +X 11 -X 4 -X 5 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 2 is the similarity parameter corresponding to the keyword, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information, X 4 is the frequency parameter corresponding to the keyword, X 5 is the word frequency parameter corresponding to non-keywords.

2. The method according to claim 1, It is characterized in that The first geographic location identifier and / or the second geographic location identifier includes an identifier corresponding to a building.

3. The method according to claim 1, It is characterized in that The obtaining of the first matching text and the second matching text comprises: obtaining the first matching text and the second matching text; Segmenting the first matching text to obtain at least one first text segmentation; Identify a first geographic location identifier from the first text segmentation; Segmenting the second matching text to obtain at least one second text segmentation; A second geographic location identifier is identified from the second text segmentation.

4. The method according to claim 3, It is characterized in that The information category includes at least one of area information, longitude and latitude information, building name information, road name information, road number information, building type information, property information, building age information, resident information, and environmental information.

5. The method according to claim 4, It is characterized in that The step of respectively determining a first word frequency of the first position participle in the first matching text and a first reverse word frequency in the sample text comprises: Using the formula Calculate the first word frequency, where tf i,1 is the first word frequency, n i,1 is the number of occurrences of the first position participle in the first matching text, ∑ k n k,1 is the number of all words in the first matching text; Using the formula Calculate the first inverse word frequency, where idf 1 is the first inverse word frequency, |D| is the number of all geographic location identifiers in the sample text, |{j:t 1 ∈d 1 }| is the number of geographic location tags containing the first position participle in the sample text.

6. The method according to claim 5, It is characterized in that The step of respectively determining a second word frequency of the second position participle in the second matching text and a second reverse word frequency in the sample text comprises: Using the formula Calculate the second word frequency, where tf i,2 is the second word frequency, n i,2 The number of occurrences of the second position participle in the second matching text, ∑ k n k,2 is the number of all words in the second matching text; Using the formula Calculate the second inverse word frequency, where idf 2 is the second inverse word frequency, |D| is the number of all geographic location identifiers in the sample text, |{j:t 2 ∈d 2 }| is the number of geographic location identifiers containing the second position participle in the sample text.

7. The method according to claim 6, It is characterized in that The calculating the similarity parameter between the first description information and the second description information based on the information category includes: In the case where the first position keyword and the second position keyword are different, the formula Y=X is used. 1 +X 6 +X 8 +X 9 +X 10 +X 11 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information.

8. The method according to claim 7, It is characterized in that The calculating the similarity parameter between the first description information and the second description information based on the information category includes: Dividing the first description information and the second description information into first information fields and second information fields corresponding to respective same information categories; Calculating a category similarity parameter corresponding to each information category according to the first information field and the second information field; The similarity parameter between the first description information and the second description information is obtained by combining the category similarity parameters.

9. The method according to claim 8, It is characterized in that The information categories respectively correspond to weight values; the similarity parameter between the first description information and the second description information is obtained by synthesizing the category similarity parameters, including: Based on the weight values ​​corresponding to the various information categories, the category similarity parameters are combined to obtain a similarity parameter between the first description information and the second description information.

10. The method according to claim 1, It is characterized in that After calculating the similarity parameter between the first description information and the second description information based on the information category, the method further includes: When the similarity parameter is not greater than the similarity threshold, it is determined that the first geographical location identifier and the second geographical location identifier are identifiers describing different geographical locations.

11. A geographical location recognition device, It is characterized in that include: A matching text acquisition module, used to acquire a first matching text and a second matching text; The first matching text contains a first geographical location identifier; The second matching text contains a second geographical location identifier; A description information search module, used to search for first description information in the first matching text, and to search for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; A similarity parameter calculation module, used to calculate a similarity parameter between the first description information and the second description information; The magnitude of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; a determination module, configured to determine, when the similarity parameter is greater than a similarity threshold, that the first geographic location identifier and the second geographic location identifier are identifiers describing the same geographic location; The first information module is used to search for first description information in the first matching text, including: segmenting the first geographical location identifier to obtain at least one first position segmentation; respectively determining a first word frequency of the first position segmentation in the first matching text and a first reverse word frequency in the sample text; The second information module is used to search for second description information in the second matching text, including: segmenting the second geographical location identifier to obtain at least one second position segmentation; and determining a second word frequency of the second position segmentation in the second matching text and a second reverse word frequency in the sample text respectively; A keyword determination module, used to determine the information category corresponding to the first description information and the second description information, including: determining a first location keyword and a first location non-keyword in the first geographic location identifier based on the first word frequency and the first reverse word frequency; determining a second location keyword and a second location non-keyword in the second geographic location identifier based on the second word frequency and the second reverse word frequency; A similarity determination module is used to calculate the similarity parameter between the first description information and the second description information based on the information category, including: when the first position keyword and the second position keyword are the same, using the formula Y=X 1 +X 2 +X 6 +X 8 +X 9 +X 10 +X 11 -X 4 -X 5 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 2 is the similarity parameter corresponding to the keyword, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information, X 4 is the frequency parameter corresponding to the keyword, X 5 is the word frequency parameter corresponding to non-keywords.

12. A geographic location identification device, comprising a memory and a processor; The memory is used to store computer program instructions; The processor is used to execute the computer program instructions to implement the following steps: obtaining a first matching text and a second matching text; the first matching text contains a first geographic location identifier; the second matching text contains a second geographic location identifier; searching for first description information in the first matching text, and searching for second description information in the second matching text; the first description information is used to describe the first geographic location identifier, and the second description information is used to describe the second geographic location identifier; calculating a similarity parameter between the first description information and the second description information ; The magnitude of the similarity parameter is used to indicate the degree of similarity between the first geographic location identifier and the second geographic location identifier; When the similarity parameter is greater than a similarity threshold, determining that the first geographical location identifier and the second geographical location identifier are identifiers describing the same geographical location; Searching for first description information in the first matching text includes: segmenting the first geographic location identifier to obtain at least one first position segmentation; determining a first word frequency of the first position segmentation in the first matching text and a first reverse word frequency in the sample text; searching for second description information in the second matching text includes: segmenting the second geographic location identifier to obtain at least one second position segmentation; determining a second word frequency of the second position segmentation in the second matching text and a second reverse word frequency in the sample text; determining information categories corresponding to the first description information and the second description information, including: determining a first position keyword and a first position non-keyword in the first geographic location identifier based on the first word frequency and the first reverse word frequency; determining a second position keyword and a second position non-keyword in the second geographic location identifier based on the second word frequency and the second reverse word frequency; calculating a similarity parameter between the first description information and the second description information based on the information category, including: when the first position keyword and the second position keyword are the same, using the formula Y=X 1 +X 2 +X 6 +X 8 +X 9 +X 10 +X 11 -X 4 -X 5 Calculate the similarity parameter, where Y is the similarity parameter, X 1 is the similarity parameter corresponding to the region information, X 2 is the similarity parameter corresponding to the keyword, X 6 is the similarity parameter corresponding to the road name information, X 8 is the similarity parameter corresponding to the property type, X 9 is the similarity parameter corresponding to the building age information, X 10 is the similarity parameter corresponding to the resident information, X 11 is the similarity parameter corresponding to the environmental information, X 4 is the frequency parameter corresponding to the keyword, X 5 is the word frequency parameter corresponding to non-keywords.

Citation Information

Patent Citations

  • A method and device for processing information

    CN109635114A