Method, apparatus and device for place name recognition based on text fragment representation learning
By constructing a place name recognition model for text fragment representation learning, using external place name entity information and prompt encoder for context encoding, combining text fragment classifiers and auxiliary tasks, the problem of language irregularity in place name recognition is solved, and the accuracy and performance of place name recognition is improved.
Patent Information
- Application Number
- CN202510456075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing deep learning models are difficult to effectively deal with the language irregularity and diversified representation methods in place name recognition, resulting in low accuracy in place name recognition.
By constructing a place name recognition model, the place name recognition task is defined as a text fragment classification task, the place name entity information is retrieved from the external knowledge database using the place name searcher, the place name entity information is retrieved from the external knowledge database, the prompt encoder is used for context representation encoding, and the text fragment representation and classification prediction are performed through the text fragment representation unit and the classifier, and the place name entity prediction auxiliary task is introduced to enhance semantic integrity.
It improves the accuracy and performance of place name recognition, enhances the understanding of the relevant context of place name, reduces noise interference, and improves the overall effect of place name recognition.
Smart Images

Figure CN119988568B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a method, apparatus, and device for place name recognition based on text fragment representation learning. Background Art
[0002] Place name recognition refers to the process of extracting geographical locations (i.e., place names) from unstructured text, which is of great significance for multiple application scenarios such as geographical information retrieval, emergency response, and natural disaster analysis. Although place name recognition is a specialized sub-task in named entity recognition, it has unique concerns. Named entity recognition generally involves identifying multiple entity types including "person", "location", and "organization", where "location" usually refers to coarse-grained place names such as countries and cities. In contrast, place name recognition not only needs to identify these large-scale place names but also extract fine-grained place name entities, including administrative units (such as countries and villages), transportation facilities (such as streets and highways), natural geographical features (such as hills and rivers), and points of interest (such as parks and schools), etc.
[0003] Currently, mainstream research mainly uses deep learning models for place name recognition, which can automatically learn the features of complex texts, thereby improving the recognition performance of the models. Common deep learning models include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs). These models model place name recognition as a sequence labeling problem, that is, classifying each word (token) to determine whether it is a place name. However, due to the inherent problems such as the ambiguity, diversity, and abbreviated forms of place names, these existing deep learning models are difficult to handle these language irregularities and diverse place name representations, limiting their effectiveness in place name recognition tasks. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, apparatus, and device for place name recognition based on text fragment representation learning to improve the accuracy of place name recognition through the effective integration of precise text fragment representations and external place name entity knowledge.
[0005] A method for place name recognition based on text fragment representation learning, the method includes:
[0006] Preprocess the text dataset and divide it into a training set, a development set, and a test set;
[0007] Construct a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text fragment classification task, aiming to identify whether each text fragment in the input text belongs to the type of place name entity; among them, the place name recognizer includes a prompt encoder, a text fragment representation unit, and a text fragment classifier;
[0008] Input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained;
[0009] Input the text to be recognized in the test set into the trained place name recognition model to predict and output the place name recognition result; among them, the place name recognition process is as follows: First, retrieve the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on a predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representations of each text segment, use the text segment classifier for classification prediction to determine whether each text segment is a place name entity.
[0010] In one embodiment, preprocess the text data set and divide it into a training set, a development set, and a test set, including:
[0011] Divide the document-level text in the data set into sentence-level text, and divide all sentence-level text into a training set, a development set, and a test set according to a preset ratio.
[0012] In one embodiment, retrieve the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever, including:
[0013] For the input text , the place name retriever uses its internal place name matching model to retrieve from the external knowledge database the set of place name entity information most relevant to the input text ; this retrieval process is based on a literal matching strategy, calculates the probability of each place name entity in the external knowledge database appearing in the input text , and ranks them according to the similarity of literal matching, and selects the top place name entities to form the external set of place name entity information ; where is the th place name entity in the set of place name entity information, and , is the total number of all place name entities in the set of place name entity information.
[0014] In one embodiment, the input text is combined with its most relevant set of geographical name entity information based on a predefined template to construct a prompt input, and a prompt encoder is used to capture the context representation of the prompt input, including:
[0015] According to the predefined template the input text is combined with its most relevant set of geographical name entity information to construct a prompt input , expressed as:
[0016] ;
[0017] wherein, represents the string concatenation operation; is a predefined template function for generating a background description of the geographical name recognition task according to the set of geographical name entity information , specifically expressed as:
[0018] ;
[0019] wherein, is a list of geographical name entities retrieved from an external knowledge database;
[0020] The prompt input is input into the prompt encoder, and the prompt encoder uses the pre-trained language model BERT to encode the prompt input to obtain the context representation of the prompt input , formally expressed as:
[0021] ;
[0022] wherein, represents the prompt encoder, represents the trainable parameters of the prompt encoder, and the context representation is the output of the last layer of the prompt encoder, including the text representation and the template representation ; wherein, and respectively represent the lengths of the input text and the template , and respectively represent the feature representations of the th character in the input text and the template .
[0023] In one embodiment, a text segment representation unit enumerates all text segments in the input text and calculates the semantic representation of each text segment based on the text representation in the context representation, including:
[0024] The text segment representation unit uses a text segment representation learning strategy to enumerate the input text all text segments composed of single words or multiple words to form a text segment set , denoted as:
[0025] ;
[0026] where, represents the number of all text segments, is the serial number, and ; the text segment , and respectively represent the starting word and the ending word in, and respectively represent the starting index and the ending index of, where and ; is a hyperparameter representing the maximum text segment length;
[0027] Based on the text representation in the context representation calculate the semantic representation of each text segment, and the semantic representation of each text segment consists of a boundary embedding and a length embedding; among them, the boundary embedding is obtained by concatenating the representations of the starting word and the ending word of the text segment, and the length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors; the text segment The semantic representation of is as follows:
[0028] ;
[0029] where, and respectively represent the feature representations of the starting word and the ending word of the text segment, represents the learned text segment length feature embedding, represents the length of the text segment;
[0030] is to capture the interaction features in the semantic representation of the text segment for text segment classification, and further input the semantic representation of the text segment into a feed-forward neural network to obtain the final representation of the text segment, denoted as:
[0031] ;
[0032] Among them, represents a feedforward neural network, represents the trainable parameters in the feedforward neural network.
[0033] In one embodiment, based on the semantic representations of the text segments, a text segment classifier is used for classification prediction to determine whether each text segment is a geographical name entity, including:
[0034] The text segment classifier converts the final representation of the text segment into a geographical name entity type score through a fully connected layer, and uses the softmax function to predict and calculate the probability distribution that the text segment belongs to the geographical name entity type, expressed as:
[0035] ;
[0036] Among them, represents the prediction result, represents the geographical name entity type, is a predefined type set and , where "geographical name" represents a geographical name entity, and "empty" represents a non-geographical name text segment; and are the trainable weights and biases respectively.
[0037] In one embodiment, the geographical name recognizer further includes a geographical name entity prediction auxiliary task unit for designing the geographical name entity prediction as an auxiliary task to learn the complete semantic representation of each geographical name entity, including:
[0038] For the predefined template , the geographical name entity information set and the template representation in the context representation of the prompt input, the geographical name entity prediction task is defined as a geographical name entity classification task. By converting the semantic representation of each geographical name entity in the geographical name entity information set into the corresponding geographical name entity type distribution, the probability distribution of each geographical name entity is obtained; specifically, first, based on the template representation calculate the semantic representation of the geographical name entity , expressed as ; Among them, and respectively represent the feature representations of the start word and the end word of and respectively represent The starting index and the ending index, represent the learned length feature embedding of the geographical name entity, indicating the length of the geographical name entity; then, the softmax function is used to calculate the probability distribution of the geographical name entity ; where represents the prediction result, represents the type of the geographical name entity, being a predefined set of types.
[0039] In one embodiment, the total training loss function of the geographical name recognition model is expressed as:
[0040] ;
[0041] ;
[0042] ;
[0043] where is the cross-entropy loss of the text segment classifier, is the cross-entropy loss of the geographical name entity prediction auxiliary task unit, and respectively represent the relative weights of each loss term; represents the number of all text segments, is the text segment, is the set of text segments, is the probability distribution that the text segment belongs to the geographical name entity type; is the set of geographical name entity information the number of all geographical name entities in, is the geographical name entity the probability distribution of; is the indicator function, if the geographical name entity type is the true label, then ; otherwise, .
[0044] A geographical name recognition device based on text segment representation learning, the device includes:
[0045] A preprocessing module, configured to preprocess a text data set and divide it into a training set, a development set, and a test set;
[0046] A model construction module, which is used to construct a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of place name entity. Among them, the place name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier;
[0047] A model training module, which is used to input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained;
[0048] A place name recognition module, which is used to input the text to be recognized in the test set into the trained place name recognition model and predict and output the place name recognition result. Among them, the place name recognition process is as follows: First, retrieve and obtain the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on a predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representation of each text segment, use the text segment classifier to perform classification prediction to determine whether each text segment is a place name entity.
[0049] A computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0050] Preprocess the text data set and divide it into a training set, a development set, and a test set;
[0051] Construct a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of place name entity. Among them, the place name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier;
[0052] Input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained;
[0053] Input the text to be recognized in the test set into the trained place name recognition model, and predict and output the place name recognition result; wherein, the place name recognition process is as follows: First, retrieve and obtain the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on the predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representations of each text segment, use the text segment classifier to perform classification prediction to determine whether each text segment is a place name entity.
[0054] The above-mentioned place name recognition method, device and equipment based on text segment representation learning have the following specific beneficial effects:
[0055] 1. By using the place name retriever to retrieve the place name entity information most relevant to the input text from the external knowledge database, it is possible to provide effective prior knowledge to guide the semantic representation learning of subsequent place name recognition, enhance the semantic representation of each text segment, and improve the accuracy of place name recognition.
[0056] 2. In the place name recognizer, the prompt encoder is used to encode the context representation of the prompt input that combines the input text and the external place name entity information, making the model's understanding of the context related to the place name in the input text more accurate; and by adopting the text segment representation learning strategy to enumerate all text segments in the input text, it is possible to more accurately capture and learn the semantic representation of each text segment in the input text, improve the accuracy of text segment classification, and thus improve the place name recognition performance.
[0057] 3. In the place name recognition model, the place name entity prediction is further introduced as an auxiliary task, which can enhance the semantic integrity of the place name entity and reduce the noise interference caused by partial overlap, thereby further strengthening the model's recognition ability for different place name entities. Brief Description of the Drawings
[0058] Figure 1 It is a schematic flowchart of the place name recognition method based on text segment representation learning in an embodiment;
[0059] Figure 2 It is a schematic diagram of the architecture of the place name recognition model in an embodiment;
[0060] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0061] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0062] In one embodiment, as Figure 1 shown, a method for place name recognition based on text fragment representation learning is provided, including the following steps:
[0063] Step S1: Preprocess the text dataset and divide it into a training set, a development set, and a test set.
[0064] Specifically, to avoid data loss problems that may occur during training and testing due to overly long document-level text lengths, the document-level text in the dataset is divided into sentence-level text, and all sentence-level text is divided into a training set, a development set, and a test set according to a preset ratio. Among them, the training set is used for iterative training of model parameters. The development set, also known as the validation set, is used for model parameter tuning and selection. The test set is used to input the model at the end to evaluate the model performance.
[0065] Step S2: Construct a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text fragment classification task, aiming to identify whether each text fragment in the input text belongs to the type of place name entity.
[0066] Among them, the architecture of the place name recognition model is as Figure 2 shown. On the one hand, the model retrieves diverse external place name entities through the place name retriever, splices the retrieved place name entity knowledge with the input text to construct a new prompt input; on the other hand, the model encodes the prompt input using a prompt encoder based on a language model, obtains a more accurate semantic representation of the text fragment through a dedicated text fragment representation unit, and performs place name recognition on each text fragment through a text fragment classifier. In addition, the model also includes a place name entity prediction auxiliary task unit for learning the complete semantic representation of the place name entity and reducing noise interference.
[0067] Given an input text , the place name recognition model needs to extract all place name entity sets from the input text, where represents the number of place name entities in the input text, represents the end word of the input text. Each place name entity is defined as a sequence of tokens , where and respectively represent 's start token and end token, and respectively represent the start and end positions of the place name entity in the input text, and satisfy .
[0068] To accurately extract all place name entities, the model regards the place name recognition task as a text segment classification task, based on a predefined type set , where "place name" represents the place name entity, and "empty" represents the text segment that is not a place name. Specifically, let be the set of all possible text segments in the input text, where represents the number of all text segments. The task of the model is to judge whether each text segment belongs to the type of "place name".
[0069] Step S3: Input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained.
[0070] Step S4: Input the text to be recognized in the test set into the trained place name recognition model, and predict and output the place name recognition result.
[0071] Specifically, the process of the place name recognition model for place name recognition includes the following steps:
[0072] 1. Place name retriever part: Retrieve and obtain the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever, including:
[0073] For the input text , the place name retriever uses its internal place name matching model to retrieve the set of place name entity information most relevant to the input text from the external knowledge database ; this retrieval process is based on the literal matching strategy, calculates the probability of each place name entity in the external knowledge database appearing in the input text , and ranks according to the similarity of literal matching, and selects the place name entities with the highest ranking to form the external set of place name entity information ; among them, is the th place name entity in the set of place name entity information, and , is the number of all place name entities in the set of place name entity information.
[0074] Specifically, in this embodiment, the noun dictionary GeoNames is selected as the external knowledge database. GeoNames provides standardized place name information from all over the world and can provide a standardized semantic representation for place name entities.
[0075] For example, for the input text “ Sheila D. Williams lives in 420 Augusta St.” ”, the place name retriever searches for matching items in the GeoNames database and calculates the relevance. Finally, the K = 3 most relevant place name entities are selected: “Augusta”, “Augusta Street” and “Augusta Springs”. Among them, the place name entity “AugustaStreet” is the standardized representation of “Augusta St.” in “420 Augusta St.” in the input text, which is convenient for identifying abbreviated place names. These place name entities provide useful prior information for the place name recognizer, thus helping the place name recognition.
[0076] 2. Place name recognizer part:
[0077] (1) Prompt encoder: Combine the input text with the set of information of its most relevant place name entities based on a predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input, including:
[0078] First, according to the predefined template combine the input text with the set of information of its most relevant place name entities to construct a prompt input , which is expressed as:
[0079] ;
[0080] Among them, represents the string concatenation operation, so that the external place name entity prompt information is combined with the original input text; is a predefined template function for generating a background description of the place name recognition task according to the set of place name entity information , which is specifically expressed as:
[0081] ;
[0082] Among them, is the list of place name entities retrieved from the external knowledge database. The prompt method based on this template can effectively guide the prompt encoder to focus on the place name-related information in the input text and enhance the recognition ability of the model. Figure 2[CLS] represents the beginning of a sentence or document, corresponding to the word vector of the first word in the prompt input. [SEP] represents the end of a sentence or document, corresponding to the word vector of the last word in the prompt input, and its function is to separate different sentences.
[0083] Then, the prompt input is input into the prompt encoder, and the prompt encoder uses the pre-trained language model BERT to encode the prompt input to obtain the context representation of the prompt input , which is formally represented as:
[0084] ;
[0085] Among them, represents the prompt encoder, represents the trainable parameters of the prompt encoder, and the context representation is the output of the last layer of the prompt encoder, including the text representation and the template representation ; among them, and respectively represent the lengths of the input text and the template , and respectively represent the feature representations of the -th character in the input text and the template .
[0086] (2) Text segment representation unit: The text segment representation unit enumerates all text segments in the input text and calculates the semantic representation of each text segment based on the text representation in the context representation, including:
[0087] To accurately capture the semantics of each word in a place name, for example, to correctly parse the abbreviated meaning of "St." in "420 Augusta St.", the text segment representation unit adopts a text segment representation learning strategy to enumerate all text segments composed of single words or multiple words in the input text to form a text segment set , and learn its semantic representation. The text segment set is represented as:
[0088] ;
[0089] Among them, represents the number of all text segments, is the serial number, and ; the text segment , and respectively represent the starting word and the ending word in and respectively represent the starting index and the ending index of and ; is a hyperparameter representing the maximum text segment length. For example, in the sentence "Paris is a romantic city", assuming the maximum text segment length L = 3, the starting and ending indices of all possible phrases are {(1,1), (2,2), (3,3), (4,4), (5,5), (6,6), (7,7), (8,8), (1,2), (2,3), (3,4), (4,5), (5,6), (6,7), (7,8), (1,3), (2,4), (3,5), (4,6), (5,7), (6,8)}. Most of these phrases are marked as "empty", except for the phrase (1,2), i.e., "Paris", which is marked as "place name".
[0090] Furthermore, based on the text representation in the context representation calculate to obtain the semantic representation of each text segment, and the semantic representation of each text segment consists of a boundary embedding and a length embedding; among them, the boundary embedding is obtained by concatenating the representations of the starting word and the ending word of the text segment, and the length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors; the text segment has the following semantic representation:
[0091] ;
[0092] wherein, and respectively represent the feature representations of the starting word and the ending word of the text segment, represents the learned text segment length feature embedding, represents the length of the text segment.
[0093] To capture the interaction features in the semantic representation of the text segment for text segment classification, further input the semantic representation of the text segment into a feed-forward neural network to obtain the final representation of the text segment, which is represented as:
[0094] ;
[0095] wherein, represents the feed-forward neural network, Represent the trainable parameters in the feedforward neural network.
[0096] (3) Text segment classifier: Based on the semantic representation of each text segment, a text segment classifier is used for classification prediction to determine whether each text segment is a geographical name entity, including:
[0097] The text segment classifier converts the final representation of the text segment into a geographical name entity type score, and uses the softmax function to predict and calculate the probability distribution that the text segment belongs to the geographical name entity type, which is expressed as:
[0098] ;
[0099] Among them, represents the prediction result, represents the geographical name entity type, is a predefined type set and , where "geographical name" represents the geographical name entity, and "empty" represents the text segment that is not a geographical name; and are the trainable weights and biases respectively.
[0100] (4) Geographical name entity prediction auxiliary task unit: In order to improve the accuracy of the semantic representation of text segments, the geographical name recognition model introduces external geographical name entity information to enrich the current task description, thereby enhancing the semantic information of the input text. Although these geographical name entity information only highlight the characteristics of geographical names through specific geographical name entities, in order to further enhance the semantic integrity of geographical name entities and reduce the noise interference caused by partial overlap, in this embodiment, a geographical name entity prediction auxiliary task unit is further added. This unit designs the geographical name entity prediction as an auxiliary task, so as to learn the complete semantic representation of each geographical name entity.
[0101] For the predefined template , the geographical name entity information set and the template representation in the context representation of the prompt input, the geographical name entity prediction task is defined as a geographical name entity classification task. By converting the semantic representation of each geographical name entity in the geographical name entity information set into the corresponding geographical name entity type distribution, the probability distribution of each geographical name entity is obtained; specifically, first, based on the template representation calculate the semantic representation of the geographical name entity , which is expressed as ; among them, and respectively represent the feature representations of the start word and end word of and respectively represent the start index and the end index of denotes the learned length feature embedding of the geographical name entity, represents the length of the geographical name entity; then, the softmax function is used to calculate the probability distribution of the geographical name entity ; where represents the prediction result, represents the type of the geographical name entity, is a predefined set of types.
[0102] Furthermore, the total training loss function of the geographical name recognition model is expressed as:
[0103] ;
[0104] ;
[0105] ;
[0106] where is the cross-entropy loss of the text segment classifier, is the cross-entropy loss of the auxiliary task unit for geographical name entity prediction, and respectively represent the relative weights of each loss term; represents the number of all text segments, is the text segment, is the set of text segments, is the probability distribution that the text segment belongs to the type of geographical name entity; is the set of geographical name entity information the number of all geographical name entities in is the geographical name entity the probability distribution of is the indicator function. If the type of the geographical name entity is the true label, then ; otherwise, .
[0107] In summary, a geographical name recognition method based on text segment representation learning provided by this application improves the performance of geographical name recognition through more accurate text segment representation learning and effective enhancement of external geographical name entity information. Furthermore, the effectiveness of this method is verified through experiments on three public datasets in this embodiment. The three open-source benchmark datasets are respectively:
[0108] GeoWebNews: This is a geotagged and geocoded dataset that contains 200 articles with a total of 6,607 place names. Among them, 2,599 place names have valid geographical coordinates, and 925 are standard place names, accounting for 35.59%. The articles are from 200 news websites and were collected from April 1st to 8th, 2018, using multilingual trigger words and topics for collection.
[0109] GeoVirus: This is a geographically parsed dataset that contains 229 articles and 2,167 annotated geographical locations, and was collected from August to September 2017.
[0110] LGL: This dataset is one of the most frequently cited geographically parsed datasets, containing 588 articles from 78 newspapers with a total of 5,088 place names, among which 3,125 (accounting for 61.47%) are standard place names. The LGL dataset mainly comes from local news, focuses on ambiguous place names, and is very suitable for evaluating the performance of place name parsing systems in specific geographical documents.
[0111] Since the document length may exceed 512 tokens (for example, 26.5% of the GeoWebNews dataset contains more than 512 tokens), this may lead to data loss during the training and testing processes. To solve this problem, in this embodiment, the NLTK tool (Natural Language Toolkit) is used to split the document-level text into multiple sentence-level texts. Then, these sentences are divided into a training set, a development set, and a test set according to the ratio of 8:1:1. Specifically:
[0112] The GeoWebNews dataset contains 1,437 sentences, among which the training set has 862, the development set has 287, and the test set has 288. The GeoVirus dataset contains 1,225 sentences, among which the training set has 735, the development set has 245, and the test set has 245. The LGL dataset contains 3,196 sentences, among which the training set has 1,917, the development set has 639, and the test set has 640.
[0113] Furthermore, precision (P), recall (R), and standard micro-averaged F1 score (F1) are further used to evaluate the performance of the place name recognition model (hereinafter referred to as RASpan) constructed in this application on three test sets, and performance comparisons are made with several existing models such as StanfordNER (Stanford Named Entity Recognition Model), Comb (Combined Model), Flair NER (Flair Named Entity Recognition Model), Flairont (Flair Ontology Enhanced Model), Stanza (Natural Language Processing Toolkit), BERT-base-NER (BERT Base Named Entity Recognition Model), SpanBERT, GEOLM (Geographical Language Model), and BERT. The evaluation and comparison results are shown in Table 1.
[0114] Table 1 Evaluation and Comparison Results
[0115]
[0116] As can be seen from Table 1, the place name recognition model constructed in this application has achieved state-of-the-art performance in the place name recognition tasks on three public data sets, verifying the effectiveness of this application.
[0117] In one embodiment, a place name recognition device based on text segment representation learning is provided, including:
[0118] A preprocessing module for preprocessing a text data set and dividing it into a training set, a development set, and a test set;
[0119] A model construction module for constructing a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of place name entity; among them, the place name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier;
[0120] A model training module for inputting the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained;
[0121] A place name recognition module for inputting the text to be recognized in the test set into the trained place name recognition model and predicting and outputting the place name recognition result; among them, the place name recognition process is as follows: First, retrieve and obtain the set of place name entity information most relevant to the input text from an external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on a predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representations of each text segment, use the text segment classifier to perform classification prediction to determine whether each text segment is a place name entity.
[0122] For the specific limitations of the place name recognition device based on text segment representation learning, reference can be made to the limitations of the place name recognition method based on text segment representation learning in the above text, which will not be elaborated here. Each module in the above-mentioned place name recognition device based on text segment representation learning can be implemented in whole or in part through software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned each module.
[0123] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structural diagram may be as shown in Figure 3 the following figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for place name recognition based on text segment representation learning. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0124] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0125] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0126] Preprocess the text data set and divide it into a training set, a development set, and a test set;
[0127] Construct a place name recognition model composed of a place name retriever and a place name recognizer. This model defines the place name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of place name entity. Among them, the place name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier;
[0128] Input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained;
[0129] Input the text to be recognized in the test set into the trained place name recognition model, and predict and output the place name recognition result; among them, the place name recognition process is as follows: First, retrieve and obtain the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on the predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representations of each text segment, use the text segment classifier to perform classification prediction to determine whether each text segment is a place name entity.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0131] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for place name recognition based on text fragment representation learning, characterized in that The method includes: Preprocessing the text dataset and dividing it into a training set, a development set, and a test set; Constructing a geographical name recognition model consisting of a geographical name retriever and a geographical name recognizer. This model defines the geographical name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of geographical name entity. Among them, the geographical name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier; Inputting the training set and the development set into the geographical name recognition model for iterative training and parameter tuning until a trained geographical name recognition model is obtained; Inputting the text to be recognized in the test set into the trained geographical name recognition model to predict and output the geographical name recognition result. Among them, the geographical name recognition process is as follows: First, retrieve and obtain the set of geographical name entity information most relevant to the input text from the external knowledge database according to the geographical name retriever; Second, combine the input text with its most relevant set of geographical name entity information based on a predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment based on the text representation in the context representation; Finally, based on the semantic representations of each text segment, use the text segment classifier to perform classification prediction to determine whether each text segment is a geographical name entity; Among them, the place name recognizer further includes a place name entity prediction auxiliary task unit, which is used to design the place name entity prediction as an auxiliary task to learn the complete semantic representation of each place name entity, including: for the predefined template , the place name entity information set and the template representation in the context representation of the prompt input , define the place name entity prediction task as a place name entity classification task, and by converting the semantic representation of each place name entity in the place name entity information set into the corresponding place name entity type distribution, obtain the probability distribution of each place name entity; specifically, first, calculate the semantic representation of the place name entity based on the template representation , expressed as ; among them, and respectively represent the feature representations of the start word and end word of , and respectively represent the start index and end index of represents the learned place name entity length feature embedding, represents the length of the place name entity; then, use the softmax function to calculate the probability distribution of the place name entity ; among them, represents the prediction result, represents the place name entity type, is a predefined type set.
2. The method according to claim 1, characterized in that Preprocessing the text dataset and dividing it into a training set, a development set, and a test set, including: Dividing the document-level text in the dataset into sentence-level text, and dividing all sentence-level text into a training set, a development set, and a test set according to a preset ratio.
3. The method according to claim 2, characterized in that, Retrieving and obtaining the set of geographical name entity information most relevant to the input text from the external knowledge database according to the geographical name retriever, including: For the input text , the place name retriever uses its internal place name matching model to retrieve a set of place name entity information most relevant to the input text from the external knowledge database ; the retrieval process is based on a literal matching strategy, calculating the probability of each place name entity in the external knowledge database appearing in the input text , and ranking them according to the similarity of literal matching, and selecting the highest-ranked place name entities to form an external set of place name entity information ; among them, is the th place name entity in the set of place name entity information, and , is the total number of all place name entities in the set of place name entity information.
4. The method according to claim 3, wherein Combining the input text with its most relevant set of geographical name entity information based on a predefined template to construct a prompt input, and using the prompt encoder to capture the context representation of the prompt input, including: According to a predefined template Combine the input text with the set of geographical name entity information most relevant to it to construct a prompt input , expressed as: ; Among them, represents a string concatenation operation; is a predefined template function for generating a background description of a place name recognition task according to the collection of place name entity information and is specifically expressed as: ; Among them, is a list of geographical name entities retrieved from an external knowledge database; The prompt input will be prompted for Input into the prompt encoder, which uses the pre-trained language model BERT to encode the prompt input and obtain the context representation of the prompt input , which is formally represented as: ; Among them, represents the prompt encoder, represents the trainable parameters of the prompt encoder, and the context representation is the output of the last layer of the prompt encoder, including the text representation and the template representation ; among them, and respectively represent the lengths of the input text and the template ; and respectively represent the feature representations of the th character in the input text and the template .
5. The method according to claim 4, wherein Using the text segment representation unit to enumerate all text segments in the input text, and calculating and obtaining the semantic representation of each text segment based on the text representation in the context representation, including: The text segment representation unit adopts a text segment representation learning strategy to enumerate all text segments composed of single words or multiple words in the input text to form a text segment set , which is expressed as: ; Among them, represents the number of all text segments, is the serial number, and ; the text segment , and respectively represent the starting word and the ending word in and respectively represent the starting index and the ending index of and ; is a hyperparameter representing the maximum text segment length; Text Representation in Context Representation Calculate and obtain the semantic representation of each text segment, where the semantic representation of each text segment consists of a boundary embedding and a length embedding; among them, the boundary embedding is obtained by concatenating the representations of the start word and the end word of the text segment, and the length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors; text segment The semantic representation is as follows: ; Among them, and respectively represent the feature representations of the start word and the end word of the text segment, represents the learned feature embedding of the text segment length, represents the length of the text segment; To capture the semantic representation of text segments and the interaction features therein for text segment classification, and further input the semantic representation of the text segment into a feedforward neural network to obtain the final representation of the text segment , which is expressed as: ; Among them, represents a feedforward neural network, represents the trainable parameters in the feedforward neural network.
6. The method according to claim 5, wherein Based on the semantic representations of each text segment, using the text segment classifier to perform classification prediction to determine whether each text segment is a geographical name entity, including: The text segment classifier converts the final representation of the text segment through a fully connected layer into a geographical entity type score, and uses the softmax function to predict and calculate the probability distribution of the text segment belonging to the geographical entity type, which is expressed as: ; Among them, represents the prediction result, represents the geographical name entity type, is a predefined type set and , where "geographical name" represents the geographical name entity, and "empty" represents the text segment that is not a geographical name; and are the trainable weights and biases respectively.
7. The method according to claim 1, wherein The total training loss function of the geographical name recognition model is expressed as: ; ; ; Among them, is the cross-entropy loss of the text segment classifier, is the cross-entropy loss of the geographical name entity prediction auxiliary task unit, and respectively represent the relative weights of each loss term; represents the number of all text segments, is a text segment, is a set of text segments, is the probability distribution that the text segment belongs to the geographical name entity type; is the set of geographical name entity information is the number of all geographical name entities in is a geographical name entity 's probability distribution; is an indicator function. If the geographical name entity type is the true label, then ; otherwise, .
8. A geographical name recognition device based on text segment representation learning, characterized in that, The device includes: A preprocessing module for preprocessing the text dataset and dividing it into a training set, a development set, and a test set; A model construction module for constructing a geographical name recognition model consisting of a geographical name retriever and a geographical name recognizer. This model defines the geographical name recognition task as a text segment classification task, aiming to identify whether each text segment in the input text belongs to the type of geographical name entity. Among them, the geographical name recognizer includes a prompt encoder, a text segment representation unit, and a text segment classifier; A model training module for inputting the training set and the development set into the geographical name recognition model for iterative training and parameter tuning until a trained geographical name recognition model is obtained; The place name recognition module is used to input the text to be recognized in the test set into the trained place name recognition model and predict and output the place name recognition result; wherein, the place name recognition process is as follows: First, retrieve and obtain the set of place name entity information most relevant to the input text from the external knowledge database according to the place name retriever; Second, combine the input text with the set of place name entity information most relevant to it based on the predefined template to construct a prompt input, and use the prompt encoder to capture the context representation of the prompt input; Then, use the text segment representation unit to enumerate all text segments in the input text, and calculate and obtain the semantic representation of each text segment according to the text representation in the context representation; Finally, based on the semantic representations of the text segments, use the text segment classifier to perform classification prediction to determine whether each text segment is a place name entity; Among them, the place name recognizer further includes a place name entity prediction auxiliary task unit, which is used to design the place name entity prediction as an auxiliary task to learn the complete semantic representation of each place name entity, including: for a predefined template , the place name entity information set and the template representation in the context representation of the prompt input , define the place name entity prediction task as a place name entity classification task. By converting the semantic representation of each place name entity in the place name entity information set into the corresponding place name entity type distribution, obtain the probability distribution of each place name entity; specifically, first, calculate the semantic representation of the place name entity based on the template representation , which is expressed as ; where and respectively represent the feature representations of the start word and end word of , and respectively represent the start index and end index of , represents the learned place name entity length feature embedding, represents the length of the place name entity; then, use the softmax function to calculate the probability distribution of the place name entity ; where represents the prediction result, represents the place name entity type, is a predefined type set. 9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Small sample named entity recognition method based on multiple tasks and prompt learning
CN116151256A