Alias mining method, POI index database construction method, recall method, system, medium and intelligent automobile
By mining the alias of POI entries from the POI entity library in the intelligent driving vehicle map software and building a POI index library, the accuracy reduction caused by deviations in user input destination entries is solved, and a higher navigation success rate is achieved.
Patent Information
- Application Number
- CN202411961914.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
In the on-board map software for intelligent driving, the destination entry entered by the user is prone to deviations due to the user's own or environmental factors, resulting in the searched alias that are far from the user's expected destination, the accuracy is reduced, and even the navigation failure is caused.
By obtaining POI entries from the POI entity library, cleaning characters according to preset cleaning rules, retaining separators, and constructing alias mining rules based on separators to dig out the alias of POI entries. Then, build a POI index library to recall alias and improve search accuracy.
By expanding the mining scope of POI alias, the accuracy of alias search is improved and the success rate of on-board navigation is ensured.
Smart Images

Figure CN119990105A_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the technical field of intelligent driving, and in particular to an alias mining method, a POI index library construction method, a recall method, a system, a medium and an intelligent car. Background Art
[0002] In conventional searches for smart driving, it is generally necessary to first obtain the user's manual or voice input terms, and then search for similar names for the user to choose. Taking the in-vehicle map software navigation as an example, in the conventional navigation process, the destination entered by the user in the in-vehicle map software is first received, and multiple similar destinations are retrieved based on the destination and displayed to the user. After the user selects and confirms, the route is displayed.
[0003] In this process, the input entries may have more words, fewer words, aliases, abbreviations, etc. due to user or environmental factors, see Table 1.
[0004] User desired destination Enter the term Deviation Type Tielu Town Railway Town Homophones Century Link Four Seasons Residence Approximate sound Jing Jia Building Liang(jing)jia Building Polyphonetic characters Shilipu 10 Lipu Digital Chinese character confusion Springfield Theatre Spring City Opera House Multi-word GT Land Plaza GT Land Few words Crooked Lake Bay Hotel Wanfuwan Hotel Typos Tsinghua Garden Community Tsinghua Garden Out of order Beijing University Peking University Abbreviation Baihe Town Baihe Town, Qingpu District Administrative district joins
[0005] Table 1 lists a variety of deviation types. Under the influence of these deviation types, the alias searched by the in-vehicle map software is far from the user's expected destination, and the accuracy is seriously reduced. In severe cases, the in-vehicle map software may even be unable to initiate normal navigation. Summary of the invention
[0006] The present specification provides an alias mining method, a POI index library construction method, a recall method, a system, a medium and a smart car to solve or partially solve the technical problem of reduced accuracy of alias search in the prior art, which can improve the accuracy of alias search and further ensure the success rate of vehicle navigation.
[0007] In order to solve the above technical problems, this specification provides an alias mining method, which includes:
[0008] Get POI entries from the POI entity library;
[0009] Cleaning the characters in the POI entry according to a preset cleaning rule, retaining all separators in the POI entry;
[0010] Extracting a target separator that meets the conditions from all separators in the POI entry;
[0011] An alias mining rule is constructed based on the target separator in the POI term, and the alias of the POI term is mined according to the alias mining rule.
[0012] This specification provides a method for constructing a POI index library, the method comprising:
[0013] Perform alias mining on POI entries in the POI entity library according to the aforementioned alias mining method to obtain POI aliases;
[0014] The POI alias is indexed and constructed according to the index construction strategy to obtain a POI index library.
[0015] This manual provides a recall method, which includes:
[0016] Extracting POI entries to be processed from the questions input by the user;
[0017] Convert the POI entry to be processed into a POI structure sequence to be processed; wherein the POI structure sequence to be processed includes one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed;
[0018] Based on the recall strategy of the POI structure sequence to be processed, recall the alias of the POI entry to be processed from the POI index library; wherein the POI index library is constructed according to the aforementioned POI index library construction method;
[0019] The alias of the POI entry to be processed is displayed to the user.
[0020] This specification provides an alias mining system, the system comprising:
[0021] An acquisition unit, used to acquire POI entries from a POI entity library;
[0022] A cleaning unit, used to clean the characters in the POI entry according to a preset cleaning rule, and retain all separators in the POI entry;
[0023] A first extraction unit, configured to extract a target separator that meets a condition from all separators in the POI entry;
[0024] The first construction unit is used to construct an alias mining rule based on the target separator in the POI term, and mine the alias of the POI term according to the alias mining rule.
[0025] This specification provides a POI index library construction system, the system comprising:
[0026] A mining unit, used to perform alias mining on POI entries in the POI entity library according to the aforementioned alias mining method to obtain POI aliases;
[0027] The second construction unit is used to perform index construction on the POI alias according to the index construction strategy to obtain a POI index library.
[0028] This specification provides a recall system, the system also includes:
[0029] A second extraction unit, used to extract the POI entries to be processed from the question input by the user;
[0030] A conversion unit, used for converting the POI entry to be processed into a POI structure sequence to be processed; wherein the POI structure sequence to be processed comprises one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed;
[0031] A recall unit, configured to recall the alias of the POI entry to be processed from a POI index library based on the recall strategy of the POI structure sequence to be processed; wherein the POI index library is constructed according to the aforementioned POI index library construction method;
[0032] A display unit is used to display the alias of the POI entry to be processed to the user.
[0033] This specification provides a smart car, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the steps of the above method are implemented when the processor executes the program.
[0034] Through one or more embodiments of this specification, this specification has the following beneficial effects or advantages:
[0035] The technical solution of this specification obtains POI entries from a POI entity library; cleans the characters in the POI entry according to a preset cleaning rule, retaining all separators in the POI entry; extracts the target separators that meet the conditions to construct an alias mining rule, and mines the aliases of the POI entry according to the alias mining rule. It can be seen that the technical solution of the present invention mines the aliases of the POI entries in the POI entity library starting from the separator aspect, expands the mining scope of the POI aliases, enables the POI entry to expand different aliases as much as possible, helps to improve the accuracy of the alias search, and thus ensures the success rate of the vehicle navigation.
[0036] The above description is only an overview of the technical solution of this specification. In order to more clearly understand the technical means of this specification, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this specification more obvious and easy to understand, the specific implementation methods of this specification are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present specification. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0038] Figure 1 A schematic diagram of a process of an alias mining method according to an embodiment of the present specification is shown;
[0039] Figure 2 A schematic diagram of a process of constructing a POI index library according to an embodiment of the present specification is shown;
[0040] Figure 3 A schematic diagram of a recall method according to an embodiment of the present specification is shown;
[0041] Figure 4 A logical flow diagram of a recall according to an embodiment of the present specification is shown;
[0042] Figure 5 to Figure 7 A schematic diagram of the overall concept flow of recall according to an embodiment of this specification is shown;
[0043] Figure 8 A schematic diagram of an alias mining system according to an embodiment of the present specification is shown;
[0044] Fig. 9 A schematic diagram of a POI index library construction system according to an embodiment of the present specification is shown;
[0045] Fig.10 A schematic diagram of a recall system according to an embodiment of the present specification is shown;
[0046] Fig.11 A schematic diagram of a smart car according to an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0047] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0048] In a first aspect of the present specification, an alias mining method is disclosed, which mainly performs alias mining on POI (Point of Interest) to expand the aliases, abbreviations, etc. of POI as much as possible.
[0049] Specifically, POI refers to any meaningful and noteworthy geographical location on the map, such as commercial places (shopping malls, restaurants, hotels), public service facilities (hospitals, schools, libraries), tourist attractions (historic sites, scenic spots), transportation hubs (train stations, bus stations, airports), etc.
[0050] A large number of POIs are collected by various means such as manual and geographic location collection vehicles, and stored in the POI entity library in the form of POI entries. POI entries contain one or more characters such as letters, numbers, words, and symbols. For example, POI entries with symbols: Wanwei Yunyan (Chengdu Merchants Big Magic Cube Store), this type of POI entry has a poor effect in mining aliases due to the separation effect of symbols.
[0051] In order to expand the POI terms into different aliases as much as possible, this manual starts from the separator to mine the aliases of the POI terms in the POI entity library. Figure 1 , the alias mining method disclosed in the present invention includes:
[0052] S101, obtaining POI entries from a POI entity library.
[0053] Specifically, the POI entity library is set locally or in the cloud, and a large number of POI entries are stored in the POI entity library. In the process of obtaining POI entries, POI entries are extracted from the POI entity library one by one or in batches.
[0054] S102, cleaning characters in the POI entry according to a preset cleaning rule, and retaining all separators in the POI entry.
[0055] The preset cleaning rules list the character types that need to be cleaned, such as Chinese, English, numbers, and blank characters.
[0056] During the cleaning process, the character types in the POI entry are scanned with reference to the character types listed in the preset cleaning rules, and the characters of the same type in the character set are removed and cleaned from the POI entry, thereby obtaining all the separators in the POI entry.
[0057] Taking the POI entry "Wanwei·Yunyan (Chengdu Merchants Big Magic Cube Store)" as an example, refer to the character types listed in the preset cleaning rules to clear all Chinese, English, numbers, and space characters in the POI entry, so that the separator of the POI entry is: ".()".
[0058] S103, extracting target delimiters that meet the conditions from all delimiters in the POI entry.
[0059] Among all the delimiters in the POI entry, there may be target delimiters that meet the conditions, and there may also be invalid delimiters that do not meet the conditions. Among them, the target delimiter is the reference delimiter for alias mining, and the invalid delimiter is not used as a reference for alias mining.
[0060] When extracting the target delimiter, the target delimiter is extracted from the delimiters in the POI entry by referring to the frequency of occurrence of the delimiter alone or referring to the entry to which the delimiter belongs alone. Of course, the target delimiter can also be extracted from the delimiters in the POI entry by combining the frequency of occurrence of the delimiter and the entry to which it belongs. The following is a detailed introduction for different reference conditions.
[0061] In the implementation process of extracting the target delimiter based on the occurrence frequency of the reference delimiter, the reference condition for extracting the target delimiter is set to: the occurrence frequency of the delimiter is greater than the set threshold M.
[0062] On this basis, the occurrence frequency of each separator among all separators in the POI entry in the POI entity library is determined; and separators with an occurrence frequency greater than a set threshold M are extracted as target separators that meet the conditions.
[0063] Assuming that the frequency of one of the two separators “.” in the POI entry is higher than the set threshold M, it is used as the target separator.
[0064] In the implementation process of extracting the target delimiter based on the semantics of the term to which the reference delimiter belongs, the reference condition for extracting the target delimiter is set to: the term to which the delimiter belongs conforms to a predefined naming rule.
[0065] Among them, the naming rules can be pre-defined manually, or the naming rules can be pre-summarized by analyzing a large number of POI entries. Further, a single separator is screened from all the separators in the POI entry. Among them, the screening method is not limited to random screening, screening according to the frequency of occurrence of the separator, screening according to the arrangement order of the separator, etc. Based on a single separator, a number of test entries containing a single separator are sampled from the POI entity library. Among them, the test entries mainly include three types: "front character + single separator", "single separator + rear character", and "front character + single separator + rear character". According to these three types, a number of test entries are sampled in the POI entity library to obtain a number of test entries. The number of test entries is judged based on the naming rules. If the number of test entries that meet the naming rules exceeds the set number threshold, the single separator is used as the target separator.
[0066] Assume that there are the following POI entries in the POI entity library: "Haide Huicheng Mall, Area 3 (Gate 7)", "Ju · Guangnian (Quanyuncun · Central Square Store)", "Kaihao Wu Seafood Grill (Qingpu Baolong Plaza Store)", "Wanwei · Yunyan (Chengdu Merchants Grand Magic Cube Store)".
[0067] Among them, all the delimiters of the POI entries are: ".()". Randomly select the delimiter ".", and sample several test entries containing "." from the POI entity library, such as "Haide ·", "Haide Huicheng Mall", "Ju · Guangnian", ". Seafood Grill", "Wanwei ·", "· Yunyan", "Wanwei · Yunyan" and other test entries. Analyze the above several test entries according to the naming rules. If more than half of the above several test entries meet the naming rules, it means that the entries divided by the delimiter "." are reasonable, then take the delimiter "." as the target delimiter.
[0068] In an optional implementation manner, in order to shorten the service time, cache the entries that contain the target delimiter and meet the naming rules for subsequent use.
[0069] It should be noted that in order to more accurately screen out the target delimiter, the above two execution logics can be combined to extract the target delimiter whose occurrence frequency in the POI entity library is greater than the set threshold and the test entries to which it belongs meet the naming rules. Of course, the entries belonging to such delimiters can also be cached for subsequent use.
[0070] S104, construct an alias mining rule based on the target delimiter in the POI entry, and mine the alias of the POI entry according to the alias mining rule.
[0071] In the process of constructing the alias mining rule, based on the target delimiter in the POI entry, obtain several reference entries including the target delimiter from the POI entity library, analyze the structural paradigm of the several reference entries, and determine the alias mining rule.
[0072] Taking the target delimiter "." as an example, reference entries such as "Haide ·", "Haide Huicheng Mall", "Ju · Guangnian", "Wanwei ·", "· Yunyan", "Wanwei · Yunyan" can be extracted from the POI entity library and their structural paradigms are analyzed, so as to construct an alias mining rule. The analyzed alias mining rules include but are not limited to:
[0073] 1. The characters before the delimiter "." are used as the POI alias.
[0074] 2. The characters after the delimiter "." are used as the POI alias.
[0075] 3. The characters before and after the delimiter "." are combined with each other as the POI alias.
[0076] Of course, in order to save service time, several reference terms including the target separation can be called from the cache for analysis, thereby constructing alias mining rules.
[0077] It is worth noting that this embodiment only takes the separator "." as an example for explanation. In practical applications, any separator that meets the above conditions can summarize its own alias mining rules, and is not limited to the separator ".".
[0078] Furthermore, the POI terms are segmented according to the alias mining rules, thereby obtaining the aliases of the POI terms.
[0079] The above is a related technical solution for mining the aliases of POI entries in the POI entity library from the separator aspect. Through the above alias mining solution, different aliases can be expanded as much as possible from the POI entries in the POI entity library, which helps to improve the accuracy of alias search, thereby ensuring the success rate of vehicle navigation.
[0080] In order to expand the scope of POI alias mining, after obtaining POI entries from the POI entity library, there are also the following alias mining methods: alias mining of POI entries based on search terms in user navigation logs, alias mining of POI entries with reference to the area where the POI is located, alias mining of POI entries with reference to the corresponding address of the POI entry, and alias mining of POI entries based on thought chains.
[0081] In the process of mining the aliases of POI terms based on the search terms in the user navigation log, a search pair is constructed based on the user navigation log. Among them, the search pair is constructed according to the structure of search term-navigation destination. For example, the search pair is: "Tianfu-Land of Abundance". Further, the search pairs in which the search term frequency is greater than the set threshold M are retained; based on the search pair, all POI terms are scanned in the POI entity library to obtain the aliases of the POI terms. Assuming that the POI term is "Land of Abundance", the POI alias of the POI term "Land of Abundance" is obtained by referring to the search pair: Tianfu.
[0082] In the process of mining aliases of POI entries with reference to the area where the POI is located, if the POI entry extracted from the POI entity library contains one or more administrative region characters from the province, city, district / county-level city / county, street / township / township / road, village / community. One or more administrative region characters are combined with each other to obtain several administrative region prefixes, and the POI entry is cleaned using the administrative region prefix to obtain the POI alias. For example, the POI entry is: "Suzhou Dushu Lake World Hotel", Suzhou City and Suzhou Dushu Lake are both administrative region prefixes, and "Suzhou City" and "Suzhou Dushu" can be combined to obtain "Suzhou City", "Suzhou Dushu", "Suzhou City Suzhou Dushu" and other administrative region prefixes. Based on this, the POI entry is cleaned and Suzhou Dushu Lake is removed. Then, the POI alias of "Suzhou Dushu Lake World Hotel" is: "Suzhou Dushu Lake World Hotel". Similarly, the POI alias of "Beijing Economic Development Zone National Xinchuang Park" is: "Xinchuang Park". Of course, all administrative region prefixes in the POI entity library may also be extracted in the above manner, and the POI entries in the POI entity library may be scanned according to all administrative region prefixes to determine the POI alias.
[0083] In the process of mining the alias of POI entries with reference to the corresponding addresses of POI entries, the corresponding addresses are found according to the POI entries. If the address contains one or more administrative region characters of province, city, district / county-level city / county, street / township / township / road, village / community, it is processed according to the process of mining the alias of POI entries in the area where the aforementioned reference POI is located, and the POI alias of the POI entry is obtained. For example, the POI entry extracted from the POI entity library is: "Zebra Smart Beijing", and the corresponding address found is: Building 8, Zone A, National Innovative Innovation Park, Beijing Economic and Technological Development Zone. According to the above public method, the administrative region prefix is removed, and the POI alias is obtained as: Building 8, Zone A, Innovative Innovation Park.
[0084] In the process of performing alias mining of POI terms based on thought chains, POI terms are obtained from a POI entity library. Based on the POI terms, task description information and task mining ideas of the POI alias mining task are constructed. K alias mining problem solving examples are obtained. Alias mining is performed on the POI terms with reference to the task description information of the POI alias mining task, the task mining ideas and the K alias mining problem solving examples to mine out the aliases of the POI terms.
[0085] Among them, the task description information includes one or more of the following: task type, task background, task purpose, output format, and storage path. Among them, the task type refers to the type of POI entry, which is helpful for the analysis of POI entry, such as information push and reminder task type, navigation and path planning task type, social interaction task type, etc. The task background is the background of the generation of the POI entry, such as the origin of the name, historical development, research and development process, etc. The task purpose includes: the predetermined number of alias mining, the word limit of alias mining and other parameters. The output format is used to limit the alias format. The storage path is used to guide the storage location of the alias.
[0086] The task mining idea is used to guide the mining of the POI terms, and describes the mining steps of the POI terms in detail, so as to guide the entire mining process. For example, the mining steps are: segmenting the POI terms, analyzing the rationality of the segmentation from the aspects of naming and semantics, and judging whether to keep the segmentation, and analyzing the reasons for keeping and removing the segmentation. Of course, different POI terms correspond to different task mining ideas.
[0087] Each alias mining problem-solving example includes: sample terms, sample term aliases, and sample problem-solving ideas. The sample terms can be POI terms that have been mined, and the sample term aliases are aliases of POI terms that have been mined. Of course, you can also construct sample terms and sample term aliases yourself. The sample problem-solving ideas refer to the overall mining logic for mining the sample term aliases from the sample terms.
[0088] The task description information of the reference POI alias mining task, the task mining ideas and the K alias mining problem-solving examples are input into the model together, so that the model mines the corresponding aliases from the POI entries according to the above information.
[0089] The above are several alias mining methods listed in this specification. The above methods can be used to fully mine the POI entity library, thereby expanding the scope of POI alias mining as much as possible.
[0090] Based on the same inventive concept, the second aspect of this specification discloses a method for constructing a POI index library. Figure 2 , the method in this specification includes:
[0091] S201, performing alias mining on POI entries in a POI entity library according to the alias mining method described in the first aspect to obtain POI aliases.
[0092] S202: index the POI aliases according to the index construction strategy to obtain a POI index library.
[0093] Of course, POI entries will also be processed in the above manner and combined with POI aliases to build a POI index library.
[0094] In the process of constructing the POI index library, the POI index library is constructed according to the word segmentation granularity and / or character length. The following describes in detail several ways of constructing the POI index library.
[0095] In the implementation method of constructing the POI index library according to the word segmentation granularity, the POI aliases are structurally divided with reference to each word segmentation granularity to obtain the structural sequence of the POI aliases at each word segmentation granularity.
[0096] The word segmentation granularity includes pinyin granularity, character granularity, and word granularity. The same POI alias is structured according to the pinyin granularity, character granularity, and word granularity.
[0097] In the division process based on the reference pinyin granularity, each character in the POI alias is converted into pinyin and marked with tones. Each pinyin is used as a division structure to obtain the corresponding structural sequence. If there are polyphones in the POI alias, the corresponding number of structural sequences is converted according to the number of polyphones. For example, the POI alias "Jingjia Building" is divided into two structural sequences: "liang / jia / da / sha" and "jing / jia / da / sha". Tones are marked in the pinyin. The symbol " / " here is only used for segmentation and has no meaning.
[0098] In the division process of the reference word granularity, each character in the POI alias is used as a division structure to obtain a corresponding structure sequence. For example, if the POI alias is "Jingjia Building", it is divided into the corresponding structure sequence: "Jing / Jia / Da / Sha".
[0099] In the process of dividing the reference word granularity, each word in the POI alias is used as a division structure to obtain a corresponding structure sequence. For example, if the POI alias is "Jingjia Building", it is divided into the corresponding structure sequence: "Jingjia / Building".
[0100] After being divided according to the above different word segmentation granularities, the same POI alias has structural sequences at different word segmentation granularities.
[0101] After obtaining the structural sequence of the POI alias at each word segmentation granularity, the pinyin granularity, the character granularity, and the word granularity are respectively used as the index directory of the POI index library. Among them, each index directory has its own index name. For example, three index directories are constructed according to the pinyin granularity, the character granularity, and the word granularity, and the index name of each index directory is: pinyin index, character index, and word index. Under each index directory, an index path of the structural sequence of the POI alias is constructed to obtain the POI index library. Specifically, the structural sequence of the POI alias is stored in each index directory according to the word segmentation granularity, and the storage location of the structural sequence of the POI alias in each index directory is used as the index path, and the structural sequence of the same granularity is located in the same index directory.
[0102] For example, the structural sequence of the POI alias "Jingjia Building" includes: "liang / jia / da / sha", "jing / jia / da / sha", "Jing / jia / Da / Sha", "Jingjia / Building". Among them, "liang / jia / da / sha", "jing / jia / da / sha" are stored under the pinyin index, and their storage positions under the pinyin index are used as the index path. "Jing / jia / Da / Sha" is stored under the character index, and its storage position under the character index is used as the index path. "Jingjia / Building" is stored under the word index, and its storage position under the word index is used as the index path.
[0103] In actual applications, a storage area is constructed locally or in the cloud as a POI index library, and the pinyin granularity, character granularity, and word granularity are used as index directories of the POI index library, respectively, and the storage area is divided into three sub-areas. Among them, the index names corresponding to each storage sub-area are: pinyin index, character index, and word index. The structural sequences of the same granularity are stored in the storage sub-area under the same index directory. For example, the structural sequence of the POI alias under the pinyin granularity is stored in the storage sub-area corresponding to the pinyin index directory, and the storage location of the structural sequence of the POI alias under the storage sub-area is used as the index path.
[0104] The POI index library constructed in the above manner has structural sequences of POI aliases in the POI index library stored in respective corresponding index directories according to word segmentation granularity for subsequent recall and use.
[0105] In the process of constructing the POI index library according to the character length, N sequence sets are set according to the character length of the POI alias; N>1 and is a positive integer. Specifically, according to the shortest character length and the longest character length in the POI alias as the boundary, one character length is divided into one sequence set, thereby obtaining N sequence sets.
[0106] Among them, different sequence sets store POI aliases of different character lengths, and the character length of the POI aliases within each sequence set remains consistent. Specifically, each sequence set corresponds to a set character length, and POI aliases that are consistent with the set character length need to be stored. Suppose there are 5 sequence sets, which store POI aliases with character lengths of 1, 2, 3, 4, and 5 respectively. If a POI alias is "Jicaoyuan" and its character length is 3, it will be stored in sequence set 3.
[0107] Different character lengths are used as index directories of the POI index library, and index names are constructed respectively. For example, the index names are: character length 1, character length 2, character length 3, character length 4, character length 5.
[0108] An index path of the POI alias is constructed under each index directory to obtain a POI index library. Specifically, the POI alias is stored in an index directory of corresponding length according to the character length, and the storage location of the POI alias is used as the starting index path.
[0109] In an optional implementation, if the N sequence sets include a target sequence set, and the number of POI aliases stored in the target sequence set is greater than the target sequence set with a set threshold, it means that the number of POI aliases stored in the target sequence set is too large. For example, if the target sequence set is sequence set 3, and the number of POI aliases stored therein is 500,000, the efficiency of subsequent recall will be reduced due to the excessive number of stored aliases, and the recall will be seriously time-consuming.
[0110] In order to improve the recall efficiency and reduce the recall time to tens of milliseconds, a sub-library storage method is adopted to split the target sequence set into M subsets, each of which stores POI aliases of the same character length, and sets the storage upper limit of each subset. Among them, M>2 and is a positive integer. Furthermore, the M subsets are named using the index name of the target sequence set, so that the index names of the M subsets are consistent with the index name of the target sequence set. Construct a separate index path for each subset, and randomly store the POI aliases in one of the M subsets according to the character length.
[0111] Continuing with the above example, the sequence set 3 is split into 3 subsets, the number of POI aliases stored in each subset is set to 200,000, and each subset is constructed with its own index path. The index names of the 3 subsets are consistent, all with a character length of 3. If the number of POI aliases with a character length of 3 is 500,000, they can be stored in three subsets, and the number of POI aliases stored in each subset is less than 200,000. In this way, the three subsets can be searched in parallel during recall, reducing service time and improving recall efficiency.
[0112] In the process of building a POI index library in combination with word segmentation granularity and character length, the POI alias is structurally divided with reference to each word segmentation granularity to obtain the structural sequence of the POI alias at each word segmentation granularity. The pinyin granularity, the character granularity, and the word granularity are respectively used as the first-level index directory of the POI index library. Under each first-level index directory, the structural sequence of POI is divided into N sequence sets according to the character length, and different character lengths are respectively used as the second-level index directories under the first-level index directory, wherein each second-level index directory has its own index name. An index path of the structural sequence of POI alias is constructed under each second-level index directory to obtain the POI index library. The structural sequence of POI alias is stored in the corresponding second-level index directory according to the character length, and the storage location of the structural sequence of POI alias under the second-level index directory is used as the index path.
[0113] Take the secondary index directory at the pinyin granularity as an example. Construct three secondary index directories, and the index names of each index directory are: pinyin length 10, pinyin length 11, and pinyin length 12. Store the structural sequence of the POI alias at the pinyin granularity in the corresponding secondary index directory according to the character length, and use the storage location of the structural sequence of the POI alias as its index path.
[0114] Of course, in order to improve the recall efficiency and reduce the recall time to tens of milliseconds, the secondary index directory can also adopt a sub-library storage method. The specific implementation method is described in the above embodiment and will not be repeated here.
[0115] The above is the technical solution for building a POI index library. By building a POI index library and optimizing the index link, the time spent on subsequent recall services can be minimized, giving users a better user experience.
[0116] Based on the same inventive concept, the third aspect of this specification discloses a recall method, which can be applied to various application scenarios such as car navigation and multimedia search. Figure 3 , the recall methods in this manual include:
[0117] S301, extracting POI entries to be processed from the question input by the user.
[0118] Specifically, the vehicle-mounted device uses a microphone in the vehicle to collect user questions input by user voice, or monitors a search box provided by the vehicle-mounted device, and can extract questions when the user manually enters them.
[0119] Considering that the POI entries to be processed may contain numeric characters, such as Section 2 of Babao Street, the numeric characters in the POI entries to be processed are changed to Chinese characters to avoid the influence of numeric characters on the accuracy of recall.
[0120] Furthermore, keywords are extracted from the user's question to obtain POI entries to be processed.
[0121] In an optional implementation, since the number of entities involved in the search task is in the order of hundreds of millions, in order to reduce time consumption, a buffer pool is set up to cache the searched POI aliases.
[0122] After extracting the POI entry to be processed, determine whether there is an alias similar to the POI entry to be processed in the buffer pool. If so, directly call the alias corresponding to the POI entry to be processed from the buffer pool and display it to the user, so as to reduce the service time as much as possible and improve the response time to provide users with a better user experience. If not, execute the solution in S302 to S304 to obtain the alias of the POI entry to be processed and display it to the user. Of course, after this, the alias of the POI entry to be processed will also be stored in the buffer pool, and the storage time will be recorded for subsequent recall and use.
[0123] S302: Convert the POI entry to be processed into a POI structure sequence to be processed.
[0124] In the specific implementation process, the POI entries to be processed are structurally divided based on one or more methods of pinyin granularity, character granularity, and word granularity, and converted into POI structure sequences to be processed. By dividing the POI entries to be processed at different word segmentation granularities, the fine granularity of the division of the POI entries to be processed is ensured to improve the accuracy of subsequent recall.
[0125] Among them, the POI structure sequence to be processed includes one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed. It is worth noting that if the POI entry to be processed has polyphones, the POI entry to be processed will be converted into different POI pinyin sequences to be processed. Suppose the POI entry to be processed is: "Jingjia Building", it is divided into two POI pinyin sequences to be processed: "liang / jia / da / sha" and "jing / jia / da / sha", each pinyin is marked with a tone, and the symbol " / " here is only used for segmentation and has no meaning.
[0126] S303: based on the recall strategy of the POI structure sequence to be processed, recall the aliases of the POI entries to be processed from the POI index library.
[0127] The POI index library is constructed according to the POI index library construction method introduced in the second aspect.
[0128] In an optional implementation, a portion of the POI index library can be preloaded into the memory with reference to the historical search log. Since the time spent reading data from the memory is one thousandth of the time spent reading data from the hard disk, the service time can be greatly reduced. For example, users often search for POI aliases such as songs and movies, which are usually concentrated in lengths of 2-10 characters. Therefore, the POI alias index of the length can be preloaded into the memory to reduce the service time.
[0129] During the recall process, see Figure 4 , comprising the following steps:
[0130] S401 , referring to the word segmentation granularity and / or character length of the POI structure sequence to be processed, extracting a plurality of POI alias structure sequences with the same word segmentation granularity and / or character length from a POI index library.
[0131] Taking the word segmentation granularity combined with character length as an example, if the two POI pinyin sequences to be processed are: "liang / jia / da / sha" and "jing / jia / da / sha", the character lengths are 13 and 12 respectively (one letter represents one character). According to the pinyin granularity + character length method, all POI alias structure sequences in the sequence set of pinyin length 12 and the sequence set of pinyin length 13 in the POI index library are selected.
[0132] S402: Determine a first alias set based on a to-be-processed POI structure sequence and a plurality of POI alias structure sequences.
[0133] Among them, each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in one or more dimensions of overall pinyin dimension, overall word dimension, overall character dimension, prefix pinyin dimension, prefix word dimension, prefix character dimension, infix pinyin dimension, infix word dimension, infix character dimension, suffix pinyin dimension, suffix word dimension, and suffix character dimension.
[0134] In the process of determining the first alias set, a similarity algorithm is used to calculate the structural similarity values of the POI structure sequence to be processed and the plurality of POI alias structure sequences, and a first alias set whose structural similarity values are greater than a similarity threshold is determined. And / or, based on the POI structure sequence to be processed, a POI alias structure sequence that is partially identical to the POI structure sequence to be processed is searched from the plurality of POI alias structure sequences to construct the first alias set.
[0135] It is worth noting that different methods may be used to determine the first alias set under different reference dimensions.
[0136] If the reference dimension is one of the overall pinyin dimension, the character dimension, and the word dimension, a similarity algorithm is used to calculate the structural similarity value of the POI structure sequence to be processed and the structural similarity value of several POI alias structure sequences. Among them, the structural similarity value is specifically one of the overall pinyin similarity value, the character similarity value, and the word similarity value. Determine the first alias set whose structural similarity value is greater than the similarity threshold.
[0137] Similarity algorithms include but are not limited to edit distance algorithms. For example, based on the edit distance algorithm, the edit distances of the POI structure sequence to be processed and several POI alias structure sequences are calculated in turn, and the first POI alias structure sequence set whose edit distance is lower than the corresponding set threshold is determined. Among them, the edit distance and the structural similarity are negatively correlated, and the shorter the edit distance, the higher the structural similarity.
[0138] If the reference dimension is one of the following: prefix pinyin dimension, prefix word dimension, prefix character dimension, infix pinyin dimension, infix word dimension, infix character dimension, suffix pinyin dimension, suffix word dimension, suffix character dimension, a POI alias structure sequence that is partially identical to the POI structure sequence to be processed is found from a plurality of POI alias structure sequences to construct a first alias set. Partial identity includes one of the following: identical prefix pinyin, identical prefix word, identical prefix character, identical infix pinyin, identical infix word, identical infix character, identical suffix pinyin, identical suffix word, identical suffix character.
[0139] In an optional implementation, in order to improve the recall accuracy, if each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in the overall pinyin dimension, after determining the first alias set based on the POI structure sequence to be processed and a plurality of POI alias structure sequences, the character similarity ratio between the POI structure sequence to be processed and each POI alias structure sequence in the first alias set is calculated; based on the character similarity ratio, a second alias set having a character similarity ratio greater than a set ratio threshold is determined from the first alias set. At this time, in the second alias set, the structural similarity value of the POI alias structure sequence is higher than the set similarity threshold and the character similarity ratio is greater than the set ratio threshold.
[0140] Taking the POI pinyin sequence "liang / jia / da / sha" to be processed as an example, the edit distance algorithm is used to determine the first POI alias structure sequence set with a distance less than or equal to 2 from the sequence set with a pinyin length of 13. Then, from the first POI alias structure sequence set, according to the formula Calculate the character similarity ratio between the pinyin sequence of the POI to be processed "liang / jia / da / sha" and each POI alias structure sequence in the set. Where S represents the character similarity ratio, Indicates the pinyin length of the POI pinyin sequence "liang / jia / da / sha" to be processed. represents the phonetic length of the i-th POI alias structure sequence in the set, Indicates the number of identical pinyins between the POI pinyin sequence "liang / jia / da / sha" to be processed and the i-th POI alias structure sequence, & indicates intersection. With the set ratio of 0.8 as the limit, determine the second POI alias structure sequence set with an edit distance less than or equal to 2 and a character similarity ratio greater than 0.8.
[0141] In an optional embodiment, in order to improve the recall accuracy, if each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in the overall word dimension or the overall character dimension, after determining the first alias set based on the POI structure sequence to be processed and a plurality of POI alias structure sequences, determine the number of identical characters between the POI structure sequence to be processed and each POI alias structure sequence in the first alias set; based on the number of identical characters, determine a third alias set from the first alias set whose number of identical characters is greater than a set character threshold. At this time, in the third alias set, the POI alias structure sequence set whose structural similarity value of the POI alias structure sequence is higher than the set similarity threshold and whose number of identical characters is greater than the set character threshold.
[0142] S403: sort the first alias set, and use one or more POI aliases that are sorted first as aliases of the POI entries to be processed.
[0143] In the process of sorting the first alias set, the importance score of each POI alias structure sequence in the first alias set can be calculated using a similarity algorithm, sorted according to the importance score, and one or more POI aliases ranked first can be selected. It is also possible to input each POI alias structure sequence in the first alias set into a sorting model, sort them using the sorting model, and select one or more POI aliases ranked first. Of course, if a second alias set or a third alias set is further selected from the first alias set, the second alias set or the third alias set can be processed similarly.
[0144] Among them, an example of a similarity algorithm is a TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, but it does not form a limitation.
[0145] Furthermore, if the reference dimension is one of the prefix pinyin dimension, prefix word dimension, prefix character dimension, infix pinyin dimension, infix word dimension, infix character dimension, suffix pinyin dimension, suffix word dimension, and suffix character dimension, the TF-IDF algorithm is used to calculate the importance score from one of the pinyin dimension, word dimension, and character dimension. If the reference dimension is the overall pinyin dimension, the TF-IDF algorithm is used to calculate the importance score from the pinyin dimension or the word dimension. If the reference dimension is the overall word dimension, the TF-IDF algorithm is used to calculate the importance score from the word dimension or the character dimension. If the reference dimension is the overall character dimension, the TF-IDF algorithm is used to calculate the importance score from the character dimension.
[0146] In the process of using the sorting model to sort the first alias set, the xgboost model and the neural network model are used as the basic model to construct the sorting model. First, the features of each POI alias structure sequence in the first alias set are extracted. Then, the sorting model is used to process the features of each POI alias structure sequence to obtain the sorting results of each POI alias structure sequence.
[0147] Among them, the extracted features include: edit distance, character features, pinyin features, recall strategy features, and word segmentation features.
[0148] Specifically, the edit distance is calculated according to the aforementioned edit distance algorithm.
[0149] The character features include: one or more of the longest common substring, the longest common subsequence, and different numbers of characters.
[0150] The pinyin features include: whether the pinyin is the same, the number of the same pinyin, the number of the same initials, the number of the same finals, and one or more of the same tones.
[0151] The recall strategy features include: recall true value, recall method field, and one or more of the recall fields. When calculating the recall method field, if the POI alias structure sequence exists in multiple recall strategies, the value of this field is the names of multiple recall strategies; if the POI alias structure sequence is recalled by only one recall strategy, the value of this field is the name of the recall strategy; otherwise, the value of this field is: not recalled. When calculating the recall field, the values of this field are: unknown, prefix pinyin, prefix word, prefix character, infix pinyin, infix word, infix character, suffix pinyin, suffix word, and suffix character. Specifically, if the prefix pinyin of the POI alias structure sequence is the pinyin of the POI entry to be processed, the value is the prefix pinyin, and others are similar.
[0152] The word segmentation features include: one or more features of the same number of word segmentations and the number of word segmentation intersections. Among them, the number of word segmentation intersections refers to the number of identical words obtained by intersecting the word segmentation intersection of the POI structure sequence to be processed and the word segmentation intersection in the POI alias structure sequence. Specifically, according to the set number of characters, for example, taking 2 characters as an example, two-character words are extracted from the POI structure sequence to be processed from left to right according to 2 characters to form a word segmentation set. Two-character words are extracted from the POI alias structure sequence from left to right according to 2 characters to form a word segmentation set. The word segmentation set of the POI structure sequence to be processed and the word segmentation set of the POI alias structure sequence are intersected, and the number of identical words obtained by the intersection is the number of word segmentation intersections.
[0153] In order to make the ranking result of POI aliases as close as possible to user expectations, this embodiment extracts features from the POI alias structure sequence from various aspects, inputs them into the ranking model to calculate the ranking score of the POI alias structure sequence. The ranking scores of each POI alias structure sequence are arranged from high to low, and the ranking result of the POI alias is obtained accordingly. One or more POI alias structure sequences ranked first are used as the aliases of the POI entries to be processed.
[0154] S304: Display the alias of the POI entry to be processed to the user.
[0155] In this embodiment, if it is a vehicle-mounted map software, and the POI entry to be processed is a geographic location, the sorting results of the aliases of the geographic location are displayed on the software interface. If the user clicks on one of the geographic location aliases, the geographic location alias is expanded and a navigation route is drawn to provide to the user. If it is a multimedia playback software, the sorting results corresponding to the multimedia aliases are displayed on the software interface. If the user clicks on one of the multimedia aliases, the song path corresponding to the multimedia alias is linked and played.
[0156] In order to further illustrate and explain the solutions in this specification, several processing methods in different scenarios are disclosed as examples below.
[0157] See also Figure 5 , is a recall solution conceived under the overall pinyin dimension. This solution is mainly suitable for solving the deviation types such as homophones, similar sounds, and polyphones that appear in users' questions. Of course, other deviation types can also be adopted.
[0158] S501, converting the POI entry to be processed from the pinyin granularity into the POI pinyin sequence to be processed. The POI entry to be processed will be extracted from the question input by the user. If the POI entry to be processed contains numbers, it can be converted into Chinese characters in advance to avoid the problem of confusion between numbers and Chinese characters.
[0159] S502: extracting a POI pinyin sequence with the same granularity as the POI pinyin sequence to be processed from the POI index library.
[0160] S503, using the edit distance algorithm to calculate the edit distance between the to-be-processed POI pinyin sequence and the POI alias pinyin sequence.
[0161] S504, calculating the character similarity ratio between the pinyin sequence of the POI to be processed and the pinyin sequence of the POI alias.
[0162] S505, determining a POI alias pinyin sequence with an edit distance less than 2 and a character similarity ratio greater than 0.8.
[0163] S506: Calculate the importance score of the POI alias pinyin sequence that meets the above conditions using the TF-IDF algorithm. Specifically, calculate the similarity score from the word dimension.
[0164] S507 , sorting the POI alias names from high to low according to the importance scores, converting the pinyin sequence of the POI alias names at the top into a Chinese sequence, and obtaining the POI alias names of the POI entries to be processed.
[0165] If the POI entry to be processed input by the user is a polyphone "Liang(jing) Jia Building", it will be converted into a pinyin sequence of the POI to be processed, and several POI alias pinyin sequences of the same granularity will be extracted from the POI index library according to the pinyin granularity. The edit distance between the two is calculated and the character similarity ratio of the pinyin is counted, so as to obtain a sequence set with an edit distance less than 2 and a similarity ratio greater than 0.8. The importance scores of each POI alias pinyin sequence in the sequence set are calculated and sorted from the pinyin dimension or word dimension using the TF-IDF algorithm, so as to obtain the POI alias pinyin sequence ranked first, which is used as an alias of "Liang(jing) Jia Building".
[0166] See also Figure 6 , is a recall solution conceived under the overall word dimension. This solution is mainly suitable for solving the types of wrong word deviation and disorder deviation that users encounter in their questions. Of course, other deviation types can also be adopted.
[0167] S601, convert the POI entry to be processed from the word granularity into the POI word sequence to be processed. The POI entry to be processed will be extracted from the question input by the user. If the POI entry to be processed contains numbers, it can be converted into Chinese characters in advance to avoid the problem of confusion between numbers and Chinese characters.
[0168] S602: extracting a POI word sequence with the same granularity as the POI word sequence to be processed from the POI index library.
[0169] S603, using an edit distance algorithm to calculate the edit distance between the POI character sequence to be processed and the POI alias name sequence.
[0170] S604, calculating the number of identical characters between the POI word sequence to be processed and the POI alias name sequence.
[0171] S605: Determine a POI alias name sequence whose edit distance is less than 2 and the number of identical characters is greater than 2.
[0172] S606: Calculate the importance score of the POI alias name sequence that meets the above conditions using the TF-IDF algorithm. Specifically, calculate the similarity score from the word dimension.
[0173] S607 , sorting the POI alias name sequence from high to low according to the importance scores, and using the POI alias name sequence at the top as the POI alias of the POI entry to be processed.
[0174] If the POI entry to be processed input by the user is "railway town", it is converted into a POI word sequence to be processed, and several POI alias name sequences with the same granularity are extracted from the POI index library according to the word granularity. The edit distance between the two is calculated and the number of identical characters is counted, so as to obtain a sequence set with an edit distance less than 2 and the number of identical characters greater than 2. The importance scores of each POI alias name sequence in the sequence set are calculated from the word dimension using the TF-IDF algorithm and sorted, so as to obtain the POI alias name sequence ranked first, which is used as the alias of "railway town".
[0175] See also Figure 7 , is a recall solution conceived in the prefix pinyin dimension, prefix word dimension, prefix character dimension, infix pinyin dimension, infix word dimension, infix character dimension, suffix pinyin dimension, suffix word dimension, and suffix character dimension. This solution is mainly suitable for solving scenarios such as missing word deviation, disorder deviation, and abbreviation in users' questions.
[0176] S701, convert the POI entry to be processed from the word granularity into the POI word sequence to be processed. The POI entry to be processed will be extracted from the question input by the user. If the POI entry to be processed contains numbers, it can be converted into Chinese characters in advance to avoid the problem of confusion between numbers and Chinese characters.
[0177] S702: extracting a POI word sequence with the same granularity as the POI word sequence to be processed from the POI index library.
[0178] S703: Find out the POI alias name sequence containing the POI character sequence to be processed from the POI alias name sequence. For example, the POI character sequence to be processed is located at the front, middle or back of the POI alias name sequence.
[0179] S704: Calculate the importance score of the POI aliases that meet the above conditions using the TF-IDF algorithm. Specifically, calculate the similarity score from the word dimension.
[0180] S705 , sorting the POI alias name sequence from high to low according to the importance scores, and using the POI alias name sequence at the top as the POI alias of the POI entry to be processed.
[0181] If the POI term to be processed input by the user is "Peking University", it is converted into a POI word sequence to be processed, and several POI alias name sequences with the same granularity are extracted from the POI index library according to the word granularity. From several POI alias name sequences, the POI alias name sequence containing "Peking University" is found to form a sequence set. The importance scores of each POI alias name sequence in the sequence set are calculated from the word dimension using the TF-IDF algorithm and sorted, so as to obtain the POI alias name sequence ranked first, which is used as the alias of "Peking University".
[0182] The above is a complete introduction to the recall solution. By recalling the aliases of the POI entries to be processed, various aliases of the POI entries to be processed can be recalled as much as possible, which can improve the accuracy of the alias search, thereby ensuring the success rate of the in-vehicle navigation.
[0183] In a fourth aspect, based on the same inventive concept as the alias mining method provided in the first aspect, the embodiment of this specification also provides an alias mining system, see Figure 8 , the system comprising:
[0184] The acquisition unit 801 is used to acquire a POI entry from a POI entity library;
[0185] A cleaning unit 802, configured to clean the characters in the POI entry according to a preset cleaning rule, and retain all separators in the POI entry;
[0186] A first extraction unit 803 is used to extract a target separator that meets the conditions from all separators in the POI entry;
[0187] The first constructing unit 804 is configured to construct an alias mining rule based on the target separator in the POI term, and mine the aliases of the POI term according to the alias mining rule.
[0188] It should be noted that the alias mining system provided by the embodiment of the present invention, wherein the specific manner in which each unit performs operations has been described in detail in the method embodiment provided in the above first aspect. The specific implementation process can refer to the method embodiment provided in the above first aspect, and will not be elaborated here.
[0189] In a fifth aspect, based on the same inventive concept as the POI index library construction method provided in the second aspect, the present specification also provides a POI index library construction system, see Fig. 9 , the system comprising:
[0190] A mining unit 901 is used to perform alias mining on POI entries in a POI entity library according to the alias mining method described in the first aspect to obtain POI aliases;
[0191] The second construction unit 902 is used to perform index construction on the POI alias according to the index construction strategy to obtain a POI index library.
[0192] It should be noted that the POI index library construction system provided by the embodiment of the present invention, wherein the specific manner in which each unit performs operations has been described in detail in the method embodiment provided in the second aspect above, and the specific implementation process can refer to the method embodiment provided in the second aspect above, and will not be elaborated here.
[0193] In a sixth aspect, based on the same inventive concept as the recall method provided in the third aspect, the present embodiment of the specification further provides a recall system, see Fig.10 , the system comprising:
[0194] The second extraction unit 1001 is used to extract the POI entries to be processed from the question input by the user;
[0195] The conversion unit 1002 is used to convert the POI entry to be processed into a POI structure sequence to be processed; wherein the POI structure sequence to be processed includes one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed;
[0196] A recall unit 1003 is used to recall the alias of the POI entry to be processed from the POI index library based on the recall strategy of the POI structure sequence to be processed; wherein the POI index library is constructed according to the POI index library construction method described in the second aspect;
[0197] The display unit 1003 is used to display the alias of the POI entry to be processed to the user.
[0198] It should be noted that the recall system provided by the embodiment of the present invention, wherein the specific manner in which each unit performs operations has been described in detail in the method embodiment provided in the above third aspect. The specific implementation process can refer to the method embodiment provided in the above third aspect, and will not be elaborated here.
[0199] Based on the same inventive concept as in the aforementioned embodiments, the seventh aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in any method embodiment in this specification.
[0200] Based on the same inventive concept as in the above-mentioned embodiment, in an eighth aspect of the present specification, a smart car is further provided. Fig.11 , including a memory 114, a processor 112, and a computer program stored in the memory 114 and executable on the processor, wherein the processor 112 implements the steps of any of the methods described above when executing the program.
[0201] Among them, Figure 3 In the embodiment of the present invention, a bus architecture (represented by bus 110) is shown, and bus 110 may include any number of interconnected buses and bridges, and bus 110 connects various circuits including one or more processors represented by processor 112 and memory represented by memory 114. Bus 110 may also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are not further described herein. Bus interface 115 provides an interface between bus 110 and receiver 111 and transmitter 113. Receiver 111 and transmitter 113 may be the same element, namely a transceiver, which provides a unit for communicating with various other devices over a transmission medium. Processor 112 is responsible for managing bus 110 and general processing, while memory 114 may be used to store data used by processor 112 when performing operations.
[0202] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general purpose systems may also be used together with the teachings based thereon. Based on the above description, it is apparent that the structure required for constructing such systems is. In addition, this specification is not directed to any particular programming language either. It should be understood that the contents of this specification described herein may be implemented using various programming languages, and the description of the specific languages above is intended to disclose the preferred implementation of this specification.
[0203] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of this description can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0204] Similarly, it should be understood that in order to streamline the present disclosure and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present specification, the various features of the present specification are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following intention: the claimed specification requires more features than the features explicitly recited in each claim. More specifically, as reflected in the claims below, the inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, with each claim itself serving as a separate embodiment of the present specification.
[0205] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0206] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of this specification and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.
[0207] The various component embodiments of this specification may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) may be used in practice to implement some or all functions of some or all components of a gateway, a proxy server, or a system according to an embodiment of this specification. This specification may also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing this specification may be stored on a computer-readable medium, or may have the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0208] It should be noted that the above embodiments illustrate rather than limit the present specification, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present specification may be implemented with the aid of hardware comprising a number of different elements and with the aid of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.
Claims
1. An alias mining method, the method comprising: Get POI entries from the POI entity library; Cleaning the characters in the POI entry according to a preset cleaning rule, retaining all separators in the POI entry; Extracting a target separator that meets the conditions from all separators in the POI entry; An alias mining rule is constructed based on the target separator in the POI term, and the alias of the POI term is mined according to the alias mining rule.
2. The method according to claim 1, wherein extracting a target separator that meets the conditions from all separators in the POI entry specifically comprises: Determine the occurrence frequency of each separator among all separators in the POI entry in the POI entity library; The separators whose occurrence frequency is greater than a set threshold are extracted as the target separators that meet the conditions.
3. The method according to claim 1 or 2, wherein the step of extracting a target separator that meets the conditions from all separators in the POI entry comprises: Filter a single delimiter from all delimiters in the POI entry; Sampling a number of test terms containing the single separator from the POI entity library; If the number of the plurality of entries to be tested that conform to the naming rule exceeds a set number threshold, the single separator is used as the target separator.
4. The method according to claim 1, after obtaining the POI entry from the POI entity library, the method further comprises: Based on the POI terms, construct task description information and task mining ideas for the POI alias mining task; wherein the task description information includes one or more of the following: task type, task background, task purpose, output format, and storage path; the task mining ideas are used to guide the mining of the POI terms; Obtain K alias mining problem-solving examples; wherein K ≥ 1 and is a positive integer; each alias mining problem-solving example includes: a sample term, a sample term alias, and a sample problem-solving idea; Aliases of the POI terms are mined by referring to the task description information of the POI alias mining task, the task mining ideas and the K alias mining problem-solving examples to mine out the aliases of the POI terms.
5. A method for constructing a POI index library, the method comprising: Perform alias mining on POI entries in a POI entity library according to the alias mining method described in any one of claims 1 to 4 to obtain POI aliases; The POI alias is indexed and constructed according to the index construction strategy to obtain a POI index library.
6. The method according to claim 5, wherein indexing the POI alias according to the index building strategy to obtain a POI index library specifically comprises: The POI alias is structurally divided with reference to each word segmentation granularity to obtain a structural sequence of the POI alias at each word segmentation granularity; wherein the word segmentation granularity is pinyin granularity, character granularity, and word granularity; The pinyin granularity, the character granularity, and the word granularity are respectively used as index directories of the POI index library, and an index path of the structural sequence of the POI alias is constructed under each of the index directories to obtain the POI index library; wherein each of the index directories has its own index name, and structural sequences of the same granularity are located under the same index directory.
7. The method according to claim 5, wherein indexing the POI alias according to the index building strategy to obtain a POI index library specifically comprises: N sequence sets are set according to the character length of the POI alias; wherein the character length in each sequence set remains consistent; N>1 and is a positive integer; Different character lengths are used as index directories of the POI index library respectively, and an index path of the POI alias is constructed under each index directory to obtain the POI index library.
8. A recall method, comprising: Extracting POI entries to be processed from the questions input by the user; Convert the POI entry to be processed into a POI structure sequence to be processed; wherein the POI structure sequence to be processed includes one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed; Based on the recall strategy of the POI structure sequence to be processed, recall the alias of the POI entry to be processed from the POI index library; wherein the POI index library is constructed according to the POI index library construction method according to any one of claims 5 to 7; The alias of the POI entry to be processed is displayed to the user.
9. The method according to claim 8, wherein the recall strategy based on the POI structure sequence to be processed, recalling the alias of the POI entry to be processed from the POI index library, specifically comprises: Referring to the word segmentation granularity and / or character length of the POI structure sequence to be processed, extracting a plurality of POI alias structure sequences with the same word segmentation granularity and / or character length from the POI index library; Based on the POI structure sequence to be processed and the several POI alias structure sequences, a first alias set is determined; wherein each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in one or more dimensions of overall pinyin dimension, overall word dimension, overall character dimension, prefix pinyin dimension, prefix word dimension, prefix character dimension, infix pinyin dimension, infix word dimension, infix character dimension, suffix pinyin dimension, suffix word dimension, and suffix character dimension. The first alias set is sorted, and one or more POI aliases that are sorted first are used as aliases of the POI entry to be processed.
10. The method according to claim 9, wherein determining the first alias set based on the to-be-processed POI structure sequence and the plurality of POI alias structure sequences specifically comprises: Calculating the structural similarity values of the to-be-processed POI structural sequence and the plurality of POI alias structural sequences by using a similarity algorithm, and determining the first alias set whose structural similarity values are greater than a similarity threshold; And / or, based on the POI structure sequence to be processed, searching for a POI alias structure sequence that is partially identical to the POI structure sequence to be processed from the plurality of POI alias structure sequences to construct the first alias set.
11. The method according to claim 9, If each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in terms of the overall pinyin dimension; After determining the first alias set based on the to-be-processed POI structure sequence and the plurality of POI alias structure sequences, the method further includes: Calculating the character similarity ratio between the to-be-processed POI structure sequence and each POI alias structure sequence in the first alias set; Based on the character similarity ratio, determining a second alias set from the first alias set whose character similarity ratio is greater than a set ratio threshold; If each POI alias structure sequence in the first alias set is similar to the POI structure sequence to be processed in terms of the overall word dimension or the overall character dimension; after determining the first alias set based on the POI structure sequence to be processed and the plurality of POI alias structure sequences, the method further includes: Determine the number of identical characters between the to-be-processed POI structure sequence and each POI alias structure sequence in the first alias set; Based on the number of identical characters, a third alias set having a number of identical characters greater than a set character threshold is determined from the first alias set.
12. An alias mining system, the system comprising: An acquisition unit, used to acquire POI entries from a POI entity library; A cleaning unit, used to clean the characters in the POI entry according to a preset cleaning rule, and retain all separators in the POI entry; A first extraction unit, configured to extract a target separator that meets a condition from all separators in the POI entry; The first construction unit is used to construct an alias mining rule based on the target separator in the POI term, and mine the alias of the POI term according to the alias mining rule.
13. A POI index library construction system, the system comprising: A mining unit, configured to perform alias mining on POI entries in a POI entity library according to the alias mining method described in any one of claims 1 to 4 to obtain POI aliases; The second construction unit is used to perform index construction on the POI alias according to the index construction strategy to obtain a POI index library.
14. A recall system, the system further comprising: A second extraction unit, used to extract the POI entries to be processed from the question input by the user; A conversion unit, used for converting the POI entry to be processed into a POI structure sequence to be processed; wherein the POI structure sequence to be processed comprises one or more of the following: a POI pinyin sequence to be processed, a POI word sequence to be processed, and a POI character sequence to be processed; A recall unit, configured to recall the alias of the POI entry to be processed from a POI index library based on the recall strategy of the POI structure sequence to be processed; wherein the POI index library is constructed according to the POI index library construction method according to any one of claims 5 to 7; A display unit is used to display the alias of the POI entry to be processed to the user.
15. An intelligent car, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 11 when executing the program.