A POI name extraction method, device, equipment and medium
By constructing a text information database for word segmentation and encoding, and combining machine learning and human experience to determine the channels, the problem of difficult POI name recognition in traditional methods is solved, and efficient extraction and real-time updating of POI names are achieved.
Patent Information
- Application Number
- CN202311059760.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-08-22
AI Technical Summary
Existing technologies struggle to effectively identify and extract POI names, especially in large corpora where traditional word segmentation methods have low matching accuracy and cannot be updated in a timely manner.
By constructing a text information database for word segmentation and encoding, calculating mutual information and adjacent character richness, and combining machine learning and human experience to identify channels, POI names are selected.
It improves the accuracy and real-time performance of POI name recognition and extraction, solves the problem of rapid changes in POI information, and ensures the validity of POI names.
Smart Images

Figure CN117077669B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for extracting POI names. Background Technology
[0002] Point of Interest (POI) is a term in Geographic Information Systems (GIS) and an important concept in the field of GIS science. It can encompass various geographically significant units with distinctive features, such as residential areas, shops, schools, and hospitals. It forms the basis for statistical summarization and analysis of relevant geographic information and is crucial for research in areas such as positioning and navigation, and regional analysis. However, obtaining the specific names of POIs is often difficult. Firstly, there is no fixed maintenance method for POIs, making it difficult to find existing POI directories or related databases when needed. Secondly, POI information changes rapidly; for example, the operating status of small shops can change frequently, making timely statistics difficult. Furthermore, there may be multiple names for the same actual location, further complicating POI name acquisition. Therefore, timely acquisition of POI names is a critical factor in improving the effectiveness of POI applications.
[0003] In practical applications of Points of Interest (POIs), large-scale corpora are often used as a basis. However, when using large-scale corpora, traditional word segmentation methods typically rely on dictionaries and forward / backward maximum matching algorithms to match existing words in the dictionary with the corpus information. But for POI name information, there is no fixed dictionary, and since POI names may not follow existing fixed words, traditional word segmentation methods cannot identify and extract POI names from the corpus information, and the obtained POI names have a low degree of matching with the corpus information. Summary of the Invention
[0004] To address the aforementioned technical problems, this specification provides one or more embodiments of a method, apparatus, device, and medium for extracting POI names.
[0005] One or more embodiments of this specification employ the following technical solutions:
[0006] This specification provides one or more embodiments of a method for extracting POI names, characterized in that the method includes:
[0007] Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information;
[0008] Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness;
[0009] The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel;
[0010] Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0011] Optionally, in one or more embodiments of this specification, a text information database is constructed by collecting batch corpus information, and word segmentation encoding is performed on each piece of corpus information in the text information database to obtain the word segmentation encoding corresponding to each piece of corpus information, specifically including:
[0012] Obtain the corpus information acquisition interface corresponding to the current application scenario, and obtain the corpus information within the preset collection period based on each of the corpus information acquisition interfaces;
[0013] The corpus information is used to write text information, and a text information database is constructed based on each piece of text information.
[0014] Determine a lexicon corresponding to the current application scenario of the text information corpus, and perform word segmentation on each of the corpus information based on the corresponding lexicon to obtain word segmentation results;
[0015] Obtain the segmented words corresponding to each corpus information in the segmentation results, obtain the matching results of the segmented words and the corresponding lexicon, and perform word encoding or character encoding on each segmented word based on the matching results to generate the segmentation code of each corpus information.
[0016] Optionally, in one or more embodiments of this specification, a preset maximum word segmentation encoding length and a preset minimum word segmentation encoding length are obtained, so as to determine the word segmentation encoding traversal range of each corpus information based on the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length;
[0017] Within the scope of the word segmentation encoding traversal, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially;
[0018] By comparing the mutual information of the word segmentation encoding with a preset mutual information threshold, and the richness of adjacent characters in the word segmentation encoding with a preset richness threshold, the candidate POI names in the corpus information are extracted.
[0019] Optionally, in one or more embodiments of this specification, the candidate POI name is input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result, specifically including:
[0020] The candidate POI name is input into the first discrimination channel, so as to transmit the candidate POI name to the corresponding client based on the first discrimination channel, and receive the first discrimination result of the corresponding client on the candidate POI name;
[0021] A machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel.
[0022] Input the candidate POI name into the second discrimination channel to output the second discrimination result.
[0023] Optionally, in one or more embodiments of this specification, before inputting the candidate POI name into a first discrimination channel to transmit the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes:
[0024] Based on the current application scenario corresponding to the batch corpus information, determine the candidate client corresponding to the current application scenario;
[0025] Determine the current working status of each of the candidate clients to identify the idle clients among the candidate clients;
[0026] Based on the historical number of judgments and historical judgment evaluations of each idle client, the processing weight corresponding to each idle client is determined, and the optimal idle client is selected as the client corresponding to the candidate POI name based on the processing weight.
[0027] Optionally, in one or more embodiments of this specification, the step of constructing a machine learning classification model based on a preset classification model structure, and training the machine learning classification model to obtain a model that meets the requirements as the second discrimination channel, specifically includes:
[0028] The word embedding structure, feature extraction structure, and classification model of the machine learning classification model are determined, so as to construct the machine learning classification model based on the word embedding structure, feature extraction structure, and classification model;
[0029] Collect POI names as positive samples and randomly extract strings of a preset length from the corresponding vocabulary as negative samples, so as to determine the training dataset of the machine learning classification model based on the positive samples and the negative samples;
[0030] The machine learning classification model is iteratively trained based on the training dataset. If the number of iterations of the iterative training is equal to a preset iteration threshold, or if the error of the machine learning classification model is less than a preset threshold, then the machine learning classification model is determined to be a model that meets the requirements, and the model that meets the requirements is used as the second discrimination channel.
[0031] Optionally, in one or more embodiments of this specification, after filtering the final POI names from the candidate POI names, the method further includes:
[0032] Obtain new word evaluation information corresponding to the final POI name;
[0033] Based on the new word evaluation information, determine the true POI name and the error POI name in the final POI name;
[0034] Determine the discrimination channel corresponding to the error POI name, and perform feedback correction on the discrimination channel.
[0035] This specification provides a POI name extraction device in one or more embodiments, the device comprising:
[0036] The word segmentation unit is used to collect batch corpus information to build a text information database, and to perform word segmentation encoding on each piece of corpus information in the text information database to obtain the word segmentation encoding corresponding to each piece of corpus information;
[0037] The acquisition unit is used to calculate the mutual information and adjacent character richness of each of the word segmentation codes, so as to obtain the candidate POI name in each of the corpus information based on the mutual information and adjacent character richness;
[0038] The discrimination unit is used to input the candidate POI names into the first discrimination channel and the second discrimination channel respectively, and obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is a machine learning classification model channel, and the second discrimination channel is a human experience channel;
[0039] The summarization unit is used to summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0040] This specification provides one or more embodiments of a POI name extraction device, comprising:
[0041] At least one processor; and,
[0042] A memory communicatively connected to the at least one processor; wherein,
[0043] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0044] Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information;
[0045] Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness;
[0046] The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel;
[0047] Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0048] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:
[0049] Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information;
[0050] Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness;
[0051] The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel;
[0052] Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0053] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0054] By segmenting and encoding the corpus information through lexical matching, and then calculating the mutual information and adjacent character richness of each segmented encoding, candidate POI names are extracted from each corpus information based on mutual information and adjacent character richness. This avoids the problem that traditional word segmentation methods cannot identify and extract POI names from the corpus information, and ensures the real-time and effectiveness of the POI name information used. By combining manual judgment or machine learning classification algorithms, it is easier to extract the final POI name from the candidate POI names based on the judgment results, thus improving the reliability of POI name extraction. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0056] Figure 1 A flowchart illustrating a method for extracting POI names provided in an embodiment of this specification;
[0057] Figure 2 A schematic diagram of the overall logic of a method for extracting POI names provided in the embodiments of this specification;
[0058] Figure 3 This is a schematic diagram of word segmentation encoding in a certain application scenario provided in the embodiments of this specification;
[0059] Figure 4 This is a schematic diagram illustrating model training in a specific application scenario provided in the embodiments of this specification.
[0060] Figure 5 A schematic diagram of the internal structure of a POI name extraction device provided for an embodiment of this specification;
[0061] Figure 6 A schematic diagram of the internal structure of a POI name extraction device provided for an embodiment of this specification;
[0062] Figure 7 This is a schematic diagram of the internal structure of a non-volatile storage medium provided in the embodiments of this specification. Detailed Implementation
[0063] This specification provides a method, apparatus, device, and medium for extracting POI names.
[0064] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0065] like Figure 1 As shown, one or more embodiments of this specification provide a method for extracting POI names, the method comprising the following steps:
[0066] S101: Collect batch corpus information to construct a text information database, and perform word segmentation encoding on each corpus information in the text information database to obtain the word segmentation encoding corresponding to each corpus information.
[0067] To analyze POI information in a corpus, extract key POI names as expected, and ensure the real-time nature and validity of the POI names used, this specification's embodiments first collect batches of corpus information to construct a text information database. Then, each piece of corpus information in the text information database is segmented and encoded to obtain its corresponding segmentation code. Specifically, in one or more embodiments of this specification, collecting batches of corpus information to construct a text information database, and then segmenting and encoding each piece of corpus information in the text information database to obtain its corresponding segmentation code, includes the following process:
[0068] First, the system obtains the corpus information acquisition interface corresponding to the current application scenario, and then acquires the corpus information within a preset collection period based on the corpus information acquisition interface. Next, the corpus information is transformed into text information, and a text information database is constructed based on each text information. In other words, in this embodiment of the specification, the corpus information is first summarized, and then each collected corpus information is treated as a text information entry to form a text information database. Taking 12345 complaint information as an example, each complaint information can be treated as a text information entry, and complaint information from a certain region and time period is collected in batches to jointly constitute the text information database.
[0069] Then, a lexicon corresponding to the current application scenario of the text information database is determined, and word segmentation results are obtained by segmenting each piece of corpus information according to the corresponding lexicon. That is, in the embodiments of this specification, word segmentation of corpus information can be performed by obtaining a lexicon of everyday language, and the lexicon can be an existing Chinese lexicon on the Internet. After obtaining the word segmentation results, the matching results between the segmented words and the corresponding lexicon are obtained according to the segmented words corresponding to each piece of corpus information in the obtained word segmentation results. Based on the matching results, word encoding or character encoding is performed on each segmented word to generate word segmentation codes for each piece of corpus information. It should also be noted that, during the encoding process, if Figure 3 As shown, each piece of corpus information is encoded in units of words according to the word segmentation results. The matched words in each piece of corpus information are represented by word codes, and the unmatched words are represented by character codes.
[0070] S102: Calculate the mutual information and adjacent character richness of each of the word segmentation codes, so as to obtain the candidate POI name in each of the corpus information based on the mutual information and adjacent character richness.
[0071] To achieve new word discovery and obtain candidate POI names from corpus information, specifically, in one or more embodiments of this specification, the mutual information and adjacent character richness of each word segmentation code are calculated to obtain candidate POI names from each corpus information based on the mutual information and adjacent character richness. This specifically includes the following steps:
[0072] First, the preset maximum and minimum word segmentation encoding lengths are obtained. Based on these lengths, the traversal range of word segmentation encodings for each corpus information is determined. Then, within this traversal range, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially. By comparing the mutual information with a preset mutual information threshold, and the adjacent character richness with a preset adjacent character richness threshold, candidate POI names are extracted from the corpus information. In addition to comparing the preset mutual information threshold and the preset adjacent character richness threshold, new words formed from the corpus information can also be extracted as candidate POI names by comparing the linear combination threshold of mutual information and adjacent character richness.
[0073] S103: Input the candidate POI names into the first discrimination channel and the second discrimination channel respectively, and obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is a machine learning classification model channel, and the second discrimination channel is a human experience channel.
[0074] In order to filter out reasonable POI names and improve the accuracy of POI name extraction, such as Figure 2In the embodiments shown in this specification, after obtaining the candidate POI names, the candidate POI names are input into the first discrimination channel of the machine learning classification module and the second discrimination channel of the human experience channel, respectively, to obtain the first discrimination result and the second discrimination result. Specifically, in one or more embodiments of this specification, inputting the candidate POI names into the first discrimination channel and the second discrimination channel to obtain the first discrimination result and the second discrimination result includes the following process:
[0075] The candidate POI names are input into the first discrimination channel, which transmits them to the corresponding client and receives the first discrimination result from the client. Then, a machine learning classification model is constructed based on a preset classification model structure. This model is trained to obtain a suitable model for the second discrimination channel. The candidate POI names are then input into the second discrimination channel to output the second discrimination result. By combining manual judgment or machine learning classification algorithms, it is beneficial to extract POI names from newly extracted words based on the judgment results. This achieves POI name extraction and solves the problem in existing technologies where the rapid change of POI information makes timely and accurate statistics difficult.
[0076] Furthermore, to improve the accuracy and speed of human judgment based on experience, in one or more embodiments of this specification, before inputting the candidate POI name into the first discrimination channel and transmitting the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes the following process:
[0077] First, based on the current application scenario corresponding to the batch corpus information, candidate clients corresponding to the current application scenario are identified. Then, the current working status of each candidate client is determined, thereby identifying the idle clients. Next, based on the historical discrimination count and historical discrimination evaluation of each idle client, the processing weight corresponding to each idle client is determined. Based on the processing weight, the optimal idle client is selected as the client corresponding to the candidate POI name.
[0078] Furthermore, in one or more embodiments of this specification, a machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel. Specifically, the process includes the following steps:
[0079] First, the word embedding structure, feature extraction structure, and classification model of the machine learning classification model are determined, and the machine learning classification model is constructed based on the word embedding structure, feature extraction structure, and classification model. That is, in a certain application scenario, the machine learning classification model can be designed using a structure of word embedding + feature extraction + classification model. The word embedding techniques that can be used include word2vector, BERT, etc., the feature extraction can use structures such as RNN, LSTM, etc., and the classification model can use logistic regression, support vector machine, XGBoost, etc.
[0080] In order to identify the selected POI names using a qualified model, this embodiment requires training a machine learning classification model. Therefore, it is necessary to collect POI names as positive samples and randomly extract strings of a preset length from the corresponding vocabulary as negative samples, so as to determine the training dataset for the machine learning classification model based on the positive and negative samples. Then, as follows... Figure 4 The machine learning classification model is iteratively trained based on the training dataset. If it is determined that the number of iterations of the current iteration training has reached the iteration threshold, or that the error of the current machine learning classification model is less than the preset threshold and meets the model conditions, then the trained machine learning classification model is used as the model that meets the requirements, and the model that meets the requirements is used as the second discrimination channel.
[0081] S104: Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0082] Based on the first and second discrimination results obtained from the above steps, the final POI names among the candidate POI names can be summarized and determined. Further, in one or more embodiments of this specification, after filtering the final POI names among the candidate POI names, in order to improve the accuracy of POI name extraction, the method further includes the following process: first, obtaining new word evaluation information corresponding to the final POI name; then, determining the true POI name and the error POI name among the final POI names based on the new word evaluation information; then, determining the discrimination channel corresponding to the error POI name, thereby providing feedback correction to the discrimination channel.
[0083] like Figure 5 As shown, one or more embodiments of this specification provide a POI name extraction device, the device comprising:
[0084] The word segmentation unit is used to collect batch corpus information to build a text information database, and to perform word segmentation encoding on each piece of corpus information in the text information database to obtain the word segmentation encoding corresponding to each piece of corpus information;
[0085] The acquisition unit is used to calculate the mutual information and adjacent character richness of each of the word segmentation codes, so as to obtain the candidate POI name in each of the corpus information based on the mutual information and adjacent character richness;
[0086] The discrimination unit is used to input the candidate POI names into the first discrimination channel and the second discrimination channel respectively, and obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is a machine learning classification model channel, and the second discrimination channel is a human experience channel;
[0087] The summarization unit is used to summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0088] like Figure 6 As shown, one or more embodiments of this specification provide a POI name extraction device, the device comprising:
[0089] At least one processor; and,
[0090] A memory communicatively connected to the at least one processor; wherein,
[0091] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0092] Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information;
[0093] Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness;
[0094] The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel;
[0095] Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0096] like Figure 7 As shown, this specification provides a schematic diagram of the internal structure of a non-volatile storage medium in one or more embodiments. Figure 7 It is known that a non-volatile storage medium stores computer-executable instructions, which are capable of:
[0097] Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information;
[0098] Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness;
[0099] The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel;
[0100] Summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names.
[0101] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0102] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0103] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for extracting POI names, characterized in that, The method includes: Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information; Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel; Summarize the first and second discrimination results to filter the final POI names from the candidate POI names; Calculate the mutual information and adjacent character richness of each of the segmentation codes, and obtain the candidate POI names from each of the corpus information based on the mutual information and adjacent character richness, specifically including: Obtain the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length, and determine the word segmentation encoding traversal range of each corpus information based on the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length; Within the scope of the word segmentation encoding traversal, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially; By comparing the mutual information of the word segmentation encoding with a preset mutual information threshold, and the richness of adjacent characters of the word segmentation encoding with a preset richness threshold, the candidate POI names in the corpus information are extracted; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively, and the first discrimination result and the second discrimination result are obtained, specifically including: The candidate POI name is input into the first discrimination channel, so as to transmit the candidate POI name to the corresponding client based on the first discrimination channel, and receive the first discrimination result of the corresponding client on the candidate POI name; A machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel. Input the candidate POI name into the second discrimination channel to output the second discrimination result; Before inputting the candidate POI name into the first discrimination channel and transmitting the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes: Based on the current application scenario corresponding to the batch corpus information, determine the candidate client corresponding to the current application scenario; Determine the current working status of each of the candidate clients to identify the idle clients among the candidate clients; Based on the historical number of judgments and historical judgment evaluations of each idle client, the processing weight corresponding to each idle client is determined, and the optimal idle client is selected as the client corresponding to the candidate POI name based on the processing weight.
2. The method for extracting POI names according to claim 1, characterized in that, Collecting batches of corpus information to construct a text information database, then performing word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information, specifically including: Obtain the corpus information acquisition interface corresponding to the current application scenario, and obtain the corpus information within the preset collection period based on each of the corpus information acquisition interfaces; The corpus information is used as text information, and a text information database is constructed based on each piece of text information. Determine a lexicon corresponding to the current application scenario of the text information corpus, and perform word segmentation on each of the corpus information based on the corresponding lexicon to obtain word segmentation results; Obtain the segmented words corresponding to each corpus information in the segmentation results, obtain the matching results of the segmented words and the corresponding lexicon, and perform word encoding or character encoding on each segmented word based on the matching results to generate the segmentation code of each corpus information.
3. The method for extracting POI names according to claim 1, characterized in that, The step of constructing a machine learning classification model based on a preset classification model structure, and training the machine learning classification model to obtain a model that meets the requirements as the second discrimination channel, specifically includes: The word embedding structure, feature extraction structure, and classification model of the machine learning classification model are determined, so as to construct the machine learning classification model based on the word embedding structure, feature extraction structure, and classification model; Collect POI names as positive samples and randomly extract strings of a preset length from the corresponding vocabulary as negative samples, so as to determine the training dataset of the machine learning classification model based on the positive samples and the negative samples; The machine learning classification model is iteratively trained based on the training dataset. If the number of iterations of the iterative training is equal to a preset iteration threshold, or if the error of the machine learning classification model is less than a preset threshold, then the machine learning classification model is determined to be a model that meets the requirements, and the model that meets the requirements is used as the second discrimination channel.
4. The method for extracting POI names according to claim 1, characterized in that, After filtering the candidate POI names to select the final POI names, the method further includes: Obtain new word evaluation information corresponding to the final POI name; Based on the new word evaluation information, determine the true POI name and the error POI name in the final POI name; Determine the discrimination channel corresponding to the error POI name, and perform feedback correction on the discrimination channel.
5. A device for extracting POI names, characterized in that, The device includes: The word segmentation unit is used to collect batch corpus information to build a text information database, and to perform word segmentation encoding on each piece of corpus information in the text information database to obtain the word segmentation encoding corresponding to each piece of corpus information; The acquisition unit is used to calculate the mutual information and adjacent character richness of each of the word segmentation codes, so as to obtain the candidate POI name in each of the corpus information based on the mutual information and adjacent character richness; The discrimination unit is used to input the candidate POI names into the first discrimination channel and the second discrimination channel respectively, and obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is a machine learning classification model channel, and the second discrimination channel is a human experience channel; The summarization unit is used to summarize the first discrimination result and the second discrimination result, and filter the final POI name from the candidate POI names; Calculate the mutual information and adjacent character richness of each of the segmentation codes, and obtain the candidate POI names from each of the corpus information based on the mutual information and adjacent character richness, specifically including: Obtain the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length, and determine the word segmentation encoding traversal range of each corpus information based on the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length; Within the scope of the word segmentation encoding traversal, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially; By comparing the mutual information of the word segmentation encoding with a preset mutual information threshold, and the richness of adjacent characters of the word segmentation encoding with a preset richness threshold, the candidate POI names in the corpus information are extracted; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively, and the first discrimination result and the second discrimination result are obtained, specifically including: The candidate POI name is input into the first discrimination channel, so as to transmit the candidate POI name to the corresponding client based on the first discrimination channel, and receive the first discrimination result of the corresponding client on the candidate POI name; A machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel. Input the candidate POI name into the second discrimination channel to output the second discrimination result; Before inputting the candidate POI name into the first discrimination channel and transmitting the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes: Based on the current application scenario corresponding to the batch corpus information, determine the candidate client corresponding to the current application scenario; Determine the current working status of each of the candidate clients to identify the idle clients among the candidate clients; Based on the historical number of judgments and historical judgment evaluations of each idle client, the processing weight corresponding to each idle client is determined, and the optimal idle client is selected as the client corresponding to the candidate POI name based on the processing weight.
6. A device for extracting POI names, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information; Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel; Summarize the first and second discrimination results to filter the final POI names from the candidate POI names; Calculate the mutual information and adjacent character richness of each of the segmentation codes, and obtain the candidate POI names from each of the corpus information based on the mutual information and adjacent character richness, specifically including: Obtain the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length, and determine the word segmentation encoding traversal range of each corpus information based on the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length; Within the scope of the word segmentation encoding traversal, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially; By comparing the mutual information of the word segmentation encoding with a preset mutual information threshold, and the richness of adjacent characters of the word segmentation encoding with a preset richness threshold, the candidate POI names in the corpus information are extracted; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively, and the first discrimination result and the second discrimination result are obtained, specifically including: The candidate POI name is input into the first discrimination channel, so as to transmit the candidate POI name to the corresponding client based on the first discrimination channel, and receive the first discrimination result of the corresponding client on the candidate POI name; A machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel. Input the candidate POI name into the second discrimination channel to output the second discrimination result; Before inputting the candidate POI name into the first discrimination channel and transmitting the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes: Based on the current application scenario corresponding to the batch corpus information, determine the candidate client corresponding to the current application scenario; Determine the current working status of each of the candidate clients to identify the idle clients among the candidate clients; Based on the historical number of judgments and historical judgment evaluations of each idle client, the processing weight corresponding to each idle client is determined, and the optimal idle client is selected as the client corresponding to the candidate POI name based on the processing weight.
7. A non-volatile storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are capable of: Collect batches of corpus information to construct a text information database, and perform word segmentation and encoding on each piece of corpus information in the text information database to obtain the word segmentation code corresponding to each piece of corpus information; Calculate the mutual information and adjacent character richness of each of the word segmentation codes, and obtain the candidate POI names in each of the corpus information based on the mutual information and adjacent character richness; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively to obtain the first discrimination result and the second discrimination result; wherein, the first discrimination channel is the machine learning classification model channel and the second discrimination channel is the human experience channel; Summarize the first and second discrimination results to filter the final POI names from the candidate POI names; Calculate the mutual information and adjacent character richness of each of the segmentation codes, and obtain the candidate POI names from each of the corpus information based on the mutual information and adjacent character richness, specifically including: Obtain the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length, and determine the word segmentation encoding traversal range of each corpus information based on the preset maximum word segmentation encoding length and the preset minimum word segmentation encoding length; Within the scope of the word segmentation encoding traversal, the mutual information and adjacent character richness of each word segmentation encoding are calculated sequentially; By comparing the mutual information of the word segmentation encoding with a preset mutual information threshold, and the richness of adjacent characters of the word segmentation encoding with a preset richness threshold, the candidate POI names in the corpus information are extracted; The candidate POI names are input into the first discrimination channel and the second discrimination channel respectively, and the first discrimination result and the second discrimination result are obtained, specifically including: The candidate POI name is input into the first discrimination channel, so as to transmit the candidate POI name to the corresponding client based on the first discrimination channel, and receive the first discrimination result of the corresponding client on the candidate POI name; A machine learning classification model is constructed based on a preset classification model structure, and the machine learning classification model is trained to obtain a model that meets the requirements as a second discrimination channel. Input the candidate POI name into the second discrimination channel to output the second discrimination result; Before inputting the candidate POI name into the first discrimination channel and transmitting the candidate POI name to the corresponding client based on the first discrimination channel, the method further includes: Based on the current application scenario corresponding to the batch corpus information, determine the candidate client corresponding to the current application scenario; Determine the current working status of each of the candidate clients to identify the idle clients among the candidate clients; Based on the historical number of judgments and historical judgment evaluations of each idle client, the processing weight corresponding to each idle client is determined, and the optimal idle client is selected as the client corresponding to the candidate POI name based on the processing weight.
Citation Information
Patent Citations
New word discovery method and system, terminal and medium
CN112966501A
Unlisted word discovery method and device, electronic equipment and storage medium
CN115034211A
Small nucleic acid drug screening method based on machine learning and expert system
CN115966260A