Place name recognition method, device and equipment based on text fragment representation learning
By adopting a text fragment representation learning method in the place name recognition model and combining external place name entity knowledge, the existing model solves the problem of place name ambiguity and diversity in place name, and achieves higher accuracy and performance in place name recognition.
Patent Information
- Application Number
- CN202510456075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing deep learning models are difficult to deal with the ambiguity, diversity and abbreviation of place names when dealing with place names recognition, resulting in limited recognition performance.
A method based on text fragment representation learning is adopted to construct a place name recognition model through the integration of precise text fragment representation and external place name entity knowledge. The model includes a place name searcher, a prompt encoder, a text fragment representation unit and a text fragment classifier. It uses an external knowledge database to retrieve relevant place name entity information, and enumerate and classify text fragments through text fragment representation learning strategies.
Through effective prior knowledge guidance and text fragment representation learning, the accuracy and performance of place name recognition are improved, the context information related to place name in the input text can be captured more accurately, and noise interference can be reduced, which improves the model's recognition ability of different place name entities.
Smart Images

Figure CN119988568A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method, device and equipment for place name recognition based on text segment representation learning. Background Art
[0002] Place name recognition refers to the process of extracting geographic locations (i.e., place names) from unstructured text, which is of great significance for multiple application scenarios such as geographic information retrieval, emergency response, and natural disaster analysis. Although place name recognition is a specialized subtask in named entity recognition, it has unique focuses. Named entity recognition usually involves identifying multiple entity types including "person", "location", and "organization", where "location" usually refers to coarse-grained place names such as countries and cities. In contrast, place name recognition not only recognizes these large-scale place names, but also needs to extract fine-grained place name entities, including administrative units (such as countries and villages), transportation facilities (such as streets and highways), natural geographical features (such as hills and rivers), and points of interest (such as parks and schools).
[0003] At present, mainstream research mainly uses deep learning models for place name recognition, which can automatically learn the features of complex texts to improve the recognition performance of the model. Common deep learning models include convolutional neural networks (CNN), recurrent neural networks (RNN) and long short-term memory networks (LSTM). These models model place name recognition as a sequence labeling problem, that is, classifying each word (token) to determine whether it is a place name. However, due to inherent problems such as the ambiguity, diversity and abbreviations of place names, these existing deep learning models have difficulty handling these linguistic irregularities and diverse place name representations, limiting their effectiveness in place name recognition tasks. Summary of the invention
[0004] Based on this, it is necessary to provide a place name recognition method, device and equipment based on text fragment representation learning to address the above technical problems, so as to improve the accuracy of place name recognition through the effective integration of precise text fragment representation and external place name entity knowledge.
[0005] A place name recognition method based on text segment representation learning, the method comprising: Preprocess the text dataset and divide it into training set, development set and test set; A place name recognition model consisting of a place name retriever and a place name recognizer is constructed. The model defines the place name recognition task as a text segment classification task, which aims to identify whether each text segment in the input text belongs to the type of place name entity. The place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier. Input the training set and development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The text to be recognized in the test set is input into the trained place name recognition model to predict and output the place name recognition result; the place name recognition process is as follows: first, the place name entity information set most relevant to the input text is retrieved from the external knowledge database according to the place name retriever; second, the input text is combined with its most relevant place name entity information set based on a predefined template to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then the text fragment representation unit is used to enumerate all the text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
[0006] In one embodiment, the text dataset is preprocessed and divided into a training set, a development set, and a test set, including: The document-level text in the dataset is divided into sentence-level text, and all sentence-level text is divided into training set, development set and test set according to preset proportions.
[0007] In one embodiment, the place name entity information set most relevant to the input text is retrieved from an external knowledge database according to the place name retriever, including: For input text The place name searcher uses its internal place name matching model to obtain the place name from the external knowledge database. Searching and entering text The most relevant set of place name entity information; the retrieval process is based on a literal matching strategy, calculating the external knowledge database Each place name entity When entering text The probability of appearing in the word is calculated and ranked according to the similarity of the literal match, and the highest ranking is selected. Place name entities form an external place name entity information collection ;in, is the first place name entities, and , It is the number of all place name entities in the place name entity information set.
[0008] In one embodiment, the input text is combined with its most relevant place name entity information set based on a predefined template to construct a prompt input, and a prompt encoder is used to capture a contextual representation of the prompt input, including: Based on predefined templates Enter text The most relevant geographical entity information collection Combine and build to get prompt input , expressed as: ; in, Represents string concatenation operation; It is a predefined template function used to collect place name entity information Generate a background description of the place name recognition task, which is specifically expressed as: ; in, is a list of place name entities retrieved from an external knowledge database; Will prompt for input Input to the prompt encoder, which uses the pre-trained language model BERT to encode the prompt input and obtain the contextual representation of the prompt input , formally expressed as: ; in, Indicates the prompt encoder, Represents the trainable parameters of the hint encoder, the context representation is the output of the last layer of the prompt encoder, containing the text representation and template representation ;in, and Respectively represent input text and templates Length, and Respectively represent input text and templates Middle Character representation.
[0009] In one embodiment, a text segment representation unit is used to enumerate all text segments in the input text, and a semantic representation of each text segment is calculated based on the text representation in the context representation, including: The text segment representation unit adopts the text segment representation learning strategy to enumerate the input text All text fragments consisting of a single word or multiple words in the text fragment set , expressed as: ; in, represents the number of all text fragments, is a sequence number, and ; Text snippet , and Respectively The starting and ending words in and Respectively The starting and ending indices of and ; is a hyperparameter indicating the maximum text segment length; Text Representation in Context-Based Representation The semantic representation of each text segment is calculated. The semantic representation of each text segment consists of boundary embedding and length embedding. The boundary embedding is obtained by connecting the representations of the start word and the end word of the text segment. The length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors. The semantic representation of is as follows: ; in, and The feature representations of the start and end words of the text segment respectively, represents the learned text segment length feature embedding, Indicates the length of the text segment; To capture the semantic representation of text fragments The interactive features in the text are used to classify text fragments and further represent the semantics of text fragments. Input into the feedforward neural network to get the final representation of the text fragment , expressed as: ; in, represents a feed-forward neural network, Represents a trainable parameter in a feedforward neural network.
[0010] In one embodiment, based on the semantic representation of each text segment, a text segment classifier is used to perform classification prediction to determine whether each text segment is a place name entity, including: The text segment classifier converts the final representation of the text segment into Converted into place name entity type score, using softmax function prediction calculation to get text fragment The probability distribution of the place name entity type is expressed as: ; in, Represents the prediction result, Indicates the place name entity type, is a predefined set of types and , where "place name" represents a place name entity and "empty" represents a text fragment that is not a place name; and are the trainable weights and biases respectively.
[0011] In one embodiment, the place name identifier further includes a place name entity prediction auxiliary task unit, which is used to design the place name entity prediction as an auxiliary task and learn the complete semantic representation of each place name entity, including: For predefined templates , Place name entity information collection and the template representation in the context representation of the prompt input , the place name entity prediction task is defined as the place name entity classification task, by collecting the place name entity information Each place name entity in The semantic representation of is converted into the corresponding place name entity type distribution, and the probability distribution of each place name entity is obtained; specifically, first, based on the template representation Counting Place Name Entities The semantic representation of ;in, and Respectively The feature representation of the starting word and the ending word of and Respectively The starting and ending indexes of represents the learned feature embedding of place name entity length, Indicates the length of the place name entity; then, the softmax function is used to calculate the place name entity The probability distribution of ;in, Represents the prediction result, Indicates the place name entity type, A predefined set of types.
[0012] In one embodiment, the total training loss function of the place name recognition model is expressed as: ; ; ; in, is the cross entropy loss of the text fragment classifier, The cross entropy loss for the auxiliary task unit for place name entity prediction, and Respectively represent the relative weight of each loss item; represents the number of all text fragments, For text snippets, is a collection of text fragments, is the probability distribution of the text segment belonging to the place name entity type; A collection of place name entity information The number of all place name entities in Place name entity The probability distribution of Is an indicator function. If the place name entity type is the true label, then ;otherwise, .
[0013] A device for identifying place names based on text segment representation learning, the device comprising: The preprocessing module is used to preprocess the text dataset and divide it into training set, development set and test set; A model building module is used to build a place name recognition model consisting of a place name retriever and a place name recognizer. The model defines the place name recognition task as a text segment classification task, the purpose of which is to identify whether each text segment in the input text belongs to the type of place name entity; wherein the place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier; The model training module is used to input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The place name recognition module is used to input the text to be recognized in the test set into the trained place name recognition model, and predict the output place name recognition results; wherein, the place name recognition process is: first, according to the place name retriever, the place name entity information set most relevant to the input text is retrieved from the external knowledge database; secondly, based on the predefined template, the input text is combined with its most relevant place name entity information set to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then, the text fragment representation unit is used to enumerate all the text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, the text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
[0014] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Preprocess the text dataset and divide it into training set, development set and test set; A place name recognition model consisting of a place name retriever and a place name recognizer is constructed. The model defines the place name recognition task as a text segment classification task, which aims to identify whether each text segment in the input text belongs to the type of place name entity. The place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier. Input the training set and development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The text to be recognized in the test set is input into the trained place name recognition model to predict and output the place name recognition result; the place name recognition process is as follows: first, the place name entity information set most relevant to the input text is retrieved from the external knowledge database according to the place name retriever; second, the input text is combined with its most relevant place name entity information set based on a predefined template to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then the text fragment representation unit is used to enumerate all the text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
[0015] The above-mentioned place name recognition method, device and apparatus based on text segment representation learning have the following beneficial effects: 1. By using a place name retriever to retrieve the place name entity information most relevant to the input text from an external knowledge database, it can provide effective prior knowledge to guide the semantic representation learning of subsequent place name recognition, enhance the semantic representation of each text fragment, and improve the accuracy of place name recognition.
[0016] 2. In the place name recognizer, the prompt encoder is used to encode the context representation of the prompt input that combines the input text and external place name entity information, so that the model has a more accurate understanding of the place name-related context in the input text; and by adopting the text fragment representation learning strategy to enumerate all the text strategies in the input text, it can more accurately capture and learn the semantic representation of each text fragment in the input text, thereby improving the accuracy of text fragment classification and thus improving the place name recognition performance.
[0017] 3. In the place name recognition model, place name entity prediction is further introduced as an auxiliary task, which can enhance the semantic integrity of place name entities and reduce noise interference caused by partial overlap, thereby further strengthening the model's recognition ability for different place name entities. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a method for place name recognition based on text segment representation learning in one embodiment; Figure 2 A schematic diagram of the architecture of a place name recognition model in one embodiment; Figure 3 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] In one embodiment, Figure 1 As shown, a place name recognition method based on text segment representation learning is provided, comprising the following steps: Step S1, preprocess the text dataset and divide it into a training set, a development set and a test set.
[0021] Specifically, in order to avoid the problem of data loss during training and testing caused by the excessive length of document-level text, the document-level text in the dataset is divided into sentence-level text, and all sentence-level text is divided into training set, development set and test set according to the preset ratio. Among them, the training set is used for iterative training of model parameters. The development set is also called the validation set, which is used for model parameter tuning and selection. The test set is used to enter the model at the end to evaluate the model performance.
[0022] Step S2, constructing a place name recognition model consisting of a place name retriever and a place name recognizer. The model defines the place name recognition task as a text segment classification task, the purpose of which is to identify whether each text segment in the input text belongs to the type of place name entity.
[0023] Among them, the place name recognition model architecture is as follows Figure 2 As shown in the figure, on the one hand, the model retrieves a variety of external place name entities through a place name retriever, and concatenates the retrieved place name entity knowledge with the input text to construct a new prompt input; on the other hand, the model encodes the prompt input using a prompt encoder based on a language model, obtains a more accurate semantic representation of the text fragment through a dedicated text fragment representation unit, and recognizes the place name of each text fragment through a text fragment classifier. In addition, the model also contains a place name entity prediction auxiliary task unit to learn a complete semantic representation of place name entities and reduce noise interference.
[0024] Given an input text In this case, the place name recognition model needs to extract all place name entity sets from the input text. ,in Indicates the number of place name entities in the input text. Indicates the end word of the input text. Each place name entity is defined as a sequence of tokens ,in and Respectively The starting and ending words of and Respectively represent the starting and ending positions of the place name entity in the input text, and satisfy .
[0025] In order to accurately extract all place name entities, the model regards the place name recognition task as a text segment classification task based on a predefined type set , where "place name" represents a place name entity, and "empty" represents a text segment that is not a place name. Specifically, let is the set of all possible text fragments in the input text, where Represents the number of all text fragments. The task of the model is to judge each text fragment Whether it belongs to the "place name" type.
[0026] Step S3, input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained.
[0027] Step S4: input the text to be recognized in the test set into the trained place name recognition model, and predict and output the place name recognition result.
[0028] Specifically, the process of place name recognition by the place name recognition model includes the following steps: 1. Place name retriever part: According to the place name retriever, the place name entity information set most relevant to the input text is retrieved from the external knowledge database, including: For input text The place name searcher uses its internal place name matching model to obtain the place name from the external knowledge database. Searching and entering text The most relevant set of place name entity information; the retrieval process is based on a literal matching strategy, calculating the external knowledge database Each place name entity When entering text The probability of appearing in the word is calculated and ranked according to the similarity of the literal match, and the highest ranking is selected. Place name entities form an external place name entity information collection ;in, is the first place name entities, and , It is the number of all place name entities in the place name entity information set.
[0029] Specifically, in this embodiment, the place name dictionary GeoNames is selected as the external knowledge database. GeoNames provides standardized place name information from all over the world and can provide standardized semantic representation for place name entities.
[0030] For example, for the input text " Sheila D. Williams lives at 420 Augusta St.” ", the place name retriever searches for matches in the GeoNames database and calculates the relevance. Finally, the most relevant K=3 place name entities are selected: "Augusta", "Augusta Street" and "Augusta Springs". Among them, the place name entity "AugustaStreet" is the normalized representation of "Augusta St." of "420 Augusta St." in the input text, which facilitates the recognition of abbreviated place names. These place name entities provide useful prior information for the place name recognizer, thereby facilitating place name recognition.
[0031] 2. Place name identifier part: (1) Prompt encoder: Based on a predefined template, the input text is combined with its most relevant place name entity information set to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input, including: First, according to the predefined template Enter text The most relevant geographical entity information collection Combine and build to get prompt input , expressed as: ; in, Indicates a string concatenation operation, which combines the external place name entity prompt information with the original input text; It is a predefined template function used to collect place name entity information Generate a background description of the place name recognition task, which is specifically expressed as: ; in, It is a list of place name entities retrieved from an external knowledge database. The prompt method based on this template can effectively guide the prompt encoder to focus on the place name related information in the input text, thus enhancing the recognition ability of the model. Figure 2 [CLS] indicates the beginning of a sentence or document, and corresponds to the word vector of the first word in the prompt input. [SEP] indicates the end of a sentence or document, and corresponds to the word vector of the last word in the prompt input. It is used to segment different sentences.
[0032] Then, enter the prompt Input to the prompt encoder, which uses the pre-trained language model BERT to encode the prompt input and obtain the contextual representation of the prompt input , formally expressed as: ; in, Indicates the prompt encoder, Represents the trainable parameters of the hint encoder, the context representation is the output of the last layer of the prompt encoder, containing the text representation and template representation ;in, and Respectively represent input text and templates Length, and Respectively represent input text and templates Middle Character representation.
[0033] (2) Text fragment representation unit: The text fragment representation unit is used to enumerate all text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation, including: In order to accurately capture the semantics of each word in a place name, for example, to correctly parse the abbreviation meaning of "St." in "420 Augusta St.", the text fragment representation unit adopts a text fragment representation learning strategy to enumerate the input text. All text fragments consisting of a single word or multiple words in the text fragment set , and learn its semantic representation. Text fragment collection It is expressed as: ; in, represents the number of all text fragments, is a sequence number, and ; Text snippet , and Respectively The starting and ending words in and Respectively The starting and ending indices of and ; is a hyperparameter representing the maximum text segment length. For example, in the sentence "Paris is a romantic city", assuming the maximum text segment length L = 3, the start and end indices of all possible phrases are {(1,1), (2,2), (3,3), (4,4), (5,5), (6,6), (7,7), (8,8), (1,2), (2,3), (3,4), (4,5), (5,6), (6,7), (7,8), (1,3), (2,4), (3,5), (4,6), (5,7), (6,8)}. Most of these phrases are marked as "empty", except for phrase (1,2), i.e. "Paris", which is marked as "place name".
[0034] Furthermore, based on the text representation in the context representation The semantic representation of each text segment is calculated. The semantic representation of each text segment consists of boundary embedding and length embedding. The boundary embedding is obtained by connecting the representations of the start word and the end word of the text segment. The length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors. The semantic representation of is as follows: ; in, and The feature representations of the start and end words of the text segment respectively, represents the learned text segment length feature embedding, Indicates the length of the text fragment.
[0035] To capture the semantic representation of text fragments The interactive features in the text are used to classify text fragments and further represent the semantics of text fragments. Input into the feedforward neural network to get the final representation of the text fragment , expressed as: ; in, represents a feed-forward neural network, Represents a trainable parameter in a feedforward neural network.
[0036] (3) Text fragment classifier: Based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity, including: The text segment classifier converts the final representation of the text segment into Converted into place name entity type score, using softmax function prediction calculation to get text fragment The probability distribution of the place name entity type is expressed as: ; in, Represents the prediction result, Indicates the place name entity type, is a predefined set of types and , where "place name" represents a place name entity and "empty" represents a text fragment that is not a place name; and are the trainable weights and biases respectively.
[0037] (4) Place name entity prediction auxiliary task unit: In order to improve the accuracy of the semantic representation of the text segment, the place name recognition model introduces external place name entity information to enrich the current task description, thereby enhancing the semantic information of the input text. Although these place name entity information only highlights the characteristics of the place name through specific place name entities, in order to further enhance the semantic integrity of the place name entity and reduce the noise interference caused by partial overlap, this embodiment further adds a place name entity prediction auxiliary task unit, which designs place name entity prediction as an auxiliary task, thereby learning the complete semantic representation of each place name entity.
[0038] For predefined templates , Place name entity information collection and the template representation in the context representation of the prompt input , the place name entity prediction task is defined as the place name entity classification task, by collecting the place name entity information Each place name entity in The semantic representation of is converted into the corresponding place name entity type distribution, and the probability distribution of each place name entity is obtained; specifically, first, based on the template representation Counting Place Name Entities The semantic representation of ;in, and Respectively The feature representation of the starting word and the ending word of and Respectively The starting and ending indexes of represents the learned feature embedding of place name entity length, Indicates the length of the place name entity; then, the softmax function is used to calculate the place name entity The probability distribution of ;in, Represents the prediction result, Indicates the place name entity type, A predefined set of types.
[0039] Furthermore, the total training loss function of the place name recognition model is expressed as: ; ; ; in, is the cross entropy loss of the text fragment classifier, The cross entropy loss for the auxiliary task unit for place name entity prediction, and Respectively represent the relative weight of each loss item; represents the number of all text fragments, For text snippets, is a collection of text fragments, is the probability distribution of the text segment belonging to the place name entity type; A collection of place name entity information The number of all place name entities in Place name entity The probability distribution of Is an indicator function. If the place name entity type is the true label, then ;otherwise, .
[0040] In summary, the present application provides a place name recognition method based on text segment representation learning, which improves the performance of place name recognition through more accurate text segment representation learning and effective external place name entity information enhancement. Furthermore, the effectiveness of the present method is verified by experiments on three public data sets in this embodiment. The three open source benchmark data sets are: GeoWebNews: This is a geo-tagged and geo-coded dataset containing 200 articles with a total of 6,607 place names. Among them, 2,599 place names have valid geographic coordinates and 925 are standard place names, accounting for 35.59%. The articles come from 200 news websites and were collected from April 1 to 8, 2018, using multilingual trigger words and themes.
[0041] GeoVirus: This is a geographically resolved dataset containing 229 articles and 2167 annotated geographic locations, collected from August to September 2017.
[0042] LGL: This dataset is one of the most commonly cited geographic resolution datasets, containing 588 articles from 78 newspapers, a total of 5088 place names, of which 3125 (61.47%) are standard place names. The LGL dataset mainly comes from local news and focuses on ambiguous place names, which is very suitable for evaluating the performance of place name resolution systems in specific geographic documents.
[0043] Since the length of a document may exceed 512 tokens (for example, 26.5% of the GeoWebNews dataset contains more than 512 tokens), this may lead to data loss during training and testing. To solve this problem, this example uses the NLTK tool (Natural Language Toolkit) to split the document-level text into multiple sentence-level texts. Then, these sentences are divided into training set, development set and test set in a ratio of 8:1:1. Specifically: The GeoWebNews dataset contains 1437 sentences, of which 862 are in the training set, 287 are in the development set, and 288 are in the test set. The GeoVirus dataset contains 1225 sentences, of which 735 are in the training set, 245 are in the development set, and 245 are in the test set. The LGL dataset contains 3196 sentences, of which 1917 are in the training set, 639 are in the development set, and 640 are in the test set.
[0044] The precision (P), recall (R) and standard micro-average F1 score (F1) were further used to evaluate the performance of the place name recognition model constructed in this application (hereinafter referred to as RASpan) on three test sets, and compared with the existing models such as StanfordNER (Stanford named entity recognition model), Comb (combination model), Flair NER (Flair named entity recognition model), Flairont (Flair ontology enhancement model), Stanza (natural language processing toolkit), BERT-base-NER (BERT-based named entity recognition model), SpanBERT, GEOLM (geographic language model) and BERT. The evaluation and comparison results are shown in Table 1.
[0045] Table 1 Evaluation and comparison results
[0046] As can be seen from Table 1, the place name recognition model constructed in this application has achieved state-of-the-art performance in the place name recognition tasks on three public datasets, verifying the effectiveness of this application.
[0047] In one embodiment, a place name recognition device based on text segment representation learning is provided, comprising: The preprocessing module is used to preprocess the text dataset and divide it into training set, development set and test set; A model building module is used to build a place name recognition model consisting of a place name retriever and a place name recognizer. The model defines the place name recognition task as a text segment classification task, the purpose of which is to identify whether each text segment in the input text belongs to the type of place name entity; wherein the place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier; The model training module is used to input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The place name recognition module is used to input the text to be recognized in the test set into the trained place name recognition model, and predict the output place name recognition results; wherein, the place name recognition process is: first, according to the place name retriever, the place name entity information set most relevant to the input text is retrieved from the external knowledge database; secondly, based on the predefined template, the input text is combined with its most relevant place name entity information set to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then, the text fragment representation unit is used to enumerate all the text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, the text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
[0048] For the specific limitations of the place name recognition device based on text segment representation learning, please refer to the limitations of the place name recognition method based on text segment representation learning above, which will not be repeated here. The various modules in the above-mentioned place name recognition device based on text segment representation learning can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0049] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a place name recognition method based on text segment representation learning is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0050] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0051] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Preprocess the text dataset and divide it into training set, development set and test set; A place name recognition model consisting of a place name retriever and a place name recognizer is constructed. The model defines the place name recognition task as a text segment classification task, which aims to identify whether each text segment in the input text belongs to the type of place name entity. The place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier. Input the training set and development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The text to be recognized in the test set is input into the trained place name recognition model to predict and output the place name recognition result; the place name recognition process is as follows: first, the place name entity information set most relevant to the input text is retrieved from the external knowledge database according to the place name retriever; second, the input text is combined with its most relevant place name entity information set based on a predefined template to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then the text fragment representation unit is used to enumerate all the text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
[0052] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0053] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements may be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A place name recognition method based on text segment representation learning, characterized in that: The method comprises: Preprocess the text dataset and divide it into training set, development set and test set; A place name recognition model consisting of a place name retriever and a place name recognizer is constructed. The model defines the place name recognition task as a text segment classification task, and the purpose is to identify whether each text segment in the input text belongs to the type of place name entity; wherein the place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier; Input the training set and development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The text to be recognized in the test set is input into the trained place name recognition model, and the place name recognition result is predicted and output; wherein, the place name recognition process is as follows: first, the place name entity information set most relevant to the input text is retrieved from the external knowledge database according to the place name retriever; second, the input text is combined with its most relevant place name entity information set based on a predefined template to construct a prompt input, and the contextual representation of the prompt input is captured using a prompt encoder; then a text fragment representation unit is used to enumerate all text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the contextual representation; finally, based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
2. The method according to claim 1, characterized in that Preprocess the text dataset and divide it into training set, development set and test set, including: The document-level text in the dataset is divided into sentence-level text, and all sentence-level text is divided into training set, development set and test set according to preset proportions.
3. The method according to claim 2, characterized in that The place name entity information set most relevant to the input text is retrieved from the external knowledge database according to the place name retriever, including: For input text The place name searcher uses its internal place name matching model to obtain the place name from the external knowledge database. Searching and entering text The most relevant set of place name entity information; the retrieval process is based on a literal matching strategy, calculating the external knowledge database Each place name entity When entering text The probability of appearing in the word is calculated and ranked according to the similarity of the literal match, and the highest ranking is selected. Place name entities form an external place name entity information collection ;in, is the first place name entities, and , It is the number of all place name entities in the place name entity information set.
4. The method according to claim 3, characterized in that Combining the input text with its most relevant place name entity information set based on a predefined template to construct a prompt input, and using a prompt encoder to capture the contextual representation of the prompt input, including: Based on predefined templates Enter text The most relevant place name entity information collection Combine and build to get prompt input , expressed as: ; in, Represents string concatenation operation; It is a predefined template function used to collect place name entity information Generate a background description of the place name recognition task, which is specifically expressed as: ; in, is a list of place name entities retrieved from an external knowledge database; Will prompt for input Input to the prompt encoder, which uses the pre-trained language model BERT to encode the prompt input to obtain the contextual representation of the prompt input , formally expressed as: ; in, Indicates the prompt encoder, Represents the trainable parameters of the hint encoder, the context representation is the output of the last layer of the prompt encoder, containing the text representation and template representation ;in, and Respectively represent input text and templates Length, and Respectively represent input text and templates Middle Character representation.
5. The method according to claim 4, characterized in that A text segment representation unit is used to enumerate all text segments in the input text, and a semantic representation of each text segment is calculated based on the text representation in the context representation, including: The text segment representation unit adopts a text segment representation learning strategy to enumerate the input text All text fragments consisting of a single word or multiple words in the text fragment set , expressed as: ; in, represents the number of all text fragments, is a sequence number, and ; Text snippet , and Respectively The starting and ending words in and Respectively The starting and ending indices of and ; is a hyperparameter indicating the maximum text segment length; Text representation based on contextual representation The semantic representation of each text segment is calculated and obtained, and the semantic representation of each text segment consists of a boundary embedding and a length embedding; wherein the boundary embedding is obtained by connecting the representations of the start word and the end word of the text segment, and the length embedding comes from a learnable lookup table that maps different text segment lengths to their respective embedding vectors; the text segment The semantic representation of is as follows: ; in, and The feature representations of the start and end words of the text segment respectively, represents the learned text segment length feature embedding, Indicates the length of the text segment; To capture the semantic representation of text fragments The interactive features in the text are used to classify text fragments and further represent the semantics of text fragments. Input into the feedforward neural network to get the final representation of the text fragment , expressed as: ; in, represents a feed-forward neural network, Represents a trainable parameter in a feedforward neural network.
6. The method according to claim 5, characterized in that Based on the semantic representation of each text segment, a text segment classifier is used to perform classification prediction to determine whether each text segment is a place name entity, including: The text segment classifier converts the final representation of the text segment into Converted into place name entity type score, using softmax function prediction calculation to get text fragment The probability distribution of the place name entity type is expressed as: ; in, Represents the prediction result, Indicates the place name entity type, is a predefined set of types and , where "place name" represents a place name entity, and "empty" represents a text fragment that is not a place name; and are the trainable weights and biases respectively.
7. The method according to claim 1, characterized in that The place name identifier further includes a place name entity prediction auxiliary task unit, which is used to design the place name entity prediction as an auxiliary task and learn the complete semantic representation of each place name entity, including: For predefined templates , Place name entity information collection and the template representation in the context representation of the prompt input , the place name entity prediction task is defined as the place name entity classification task, by collecting the place name entity information Each place name entity in The semantic representation of is converted into the corresponding place name entity type distribution, and the probability distribution of each place name entity is obtained; specifically, first, based on the template representation Counting Place Name Entities The semantic representation of ;in, and Respectively The feature representation of the starting word and the ending word of and Respectively The starting and ending indexes of represents the learned feature embedding of place name entity length, Indicates the length of the place name entity; then, the softmax function is used to calculate the place name entity The probability distribution of ;in, Represents the prediction result, Indicates the place name entity type, A predefined set of types.
8. The method according to claim 7, characterized in that The total training loss function of the place name recognition model is expressed as: ; ; ; in, is the cross entropy loss of the text fragment classifier, The cross entropy loss for the auxiliary task unit for place name entity prediction, and Respectively represent the relative weight of each loss item; represents the number of all text fragments, For text snippets, is a collection of text fragments, is the probability distribution of the text segment belonging to the place name entity type; A collection of place name entity information The number of all place name entities in Place name entity The probability distribution of Is an indicator function. If the place name entity type is the true label, then ;otherwise, .
9. A place name recognition device based on text segment representation learning, characterized in that: The device comprises: The preprocessing module is used to preprocess the text dataset and divide it into training set, development set and test set; A model building module is used to build a place name recognition model consisting of a place name retriever and a place name recognizer. The model defines the place name recognition task as a text segment classification task, the purpose of which is to identify whether each text segment in the input text belongs to the type of place name entity; wherein the place name recognizer includes a prompt encoder, a text segment representation unit and a text segment classifier; The model training module is used to input the training set and the development set into the place name recognition model for iterative training and parameter tuning until a trained place name recognition model is obtained; The place name recognition module is used to input the text to be recognized in the test set into the trained place name recognition model, and predict the output place name recognition result; wherein, the place name recognition process is: first, according to the place name retriever, the place name entity information set most relevant to the input text is retrieved from the external knowledge database; secondly, based on the predefined template, the input text is combined with its most relevant place name entity information set to construct a prompt input, and the prompt encoder is used to capture the contextual representation of the prompt input; then, a text fragment representation unit is used to enumerate all text fragments in the input text, and the semantic representation of each text fragment is calculated based on the text representation in the context representation; finally, based on the semantic representation of each text fragment, a text fragment classifier is used to perform classification prediction to determine whether each text fragment is a place name entity.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Chinese named entity recognition method and device for dynamically fusing dictionary information
CN113988074A
Small sample named entity recognition method based on multiple tasks and prompt learning
CN116151256A
Two-stage named entity recognition method based on prompt learning
CN117236335A
Ancient language named entity recognition method based on knowledge embedding
CN117272998A
AU2020103654A4