Named entity extraction method based on lexical item combination

Through a method based on term merging, named entities are automatically identified, which solves the problems of Chinese proper noun segmentation and long entity recognition, and realizes efficient and low-resource named entity extraction, which is suitable for a variety of application scenarios.

CN120805912AActive Publication Date: 2025-10-17CENT SOUTH UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511303886.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing named entity recognition methods have problems with mis-segmentation and incomplete recognition in Chinese proper noun segmentation and long entity recognition. They also rely on external models and dictionaries, which have high maintenance costs and are difficult to adapt to fields with frequent semantic changes and low-resource application scenarios.

Method used

A term merging-based method is used to automatically identify named entities through term combinations and frequency statistics after word segmentation. This includes sequential combination judgment, full-text frequency statistics, and term merging, thereby achieving named entity extraction without the need for external labeled data or pre-trained models.

Benefits of technology

It achieves efficient recognition of named entities without relying on external dictionaries or models. It is suitable for open domains and low-resource scenarios, reduces computing resource consumption, and is suitable for lightweight applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805912A_ABST
    Figure CN120805912A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text lexical item recognition, and particularly discloses a lexical item merging-based named entity extraction method, which comprises the following steps of: sequentially performing sequential combination judgment, full-text frequency statistics and lexical item merging, and extracting a named entity on the premise of not depending on an external dictionary or a training model. According to the method, continuous and high-frequency lexical item combinations can be automatically mined and merged from original texts, named entity extraction is achieved, the method is suitable for open domain, low-resource or real-time text processing scenes, model loading or training is not needed, and the method can efficiently run on terminals or edge devices with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text term recognition, and specifically discloses a method for extracting named entities based on term merging. Background Art

[0002] Named Entity Recognition (NER) is one of the core tasks in natural language processing. Its recognition accuracy directly affects downstream applications such as text structuring, information extraction, and knowledge graph construction.

[0003] Most existing named entity recognition methods are based on machine learning or deep learning technologies. These methods typically rely on manually annotated corpora, large-scale pre-trained language models, or externally imported, manually constructed entity dictionaries for training and inference. In particular, these methods require large-scale pre-trained language models or manually constructed entity dictionaries as prior knowledge. Although these methods demonstrate high accuracy on standard datasets, they still face several technical challenges in practical applications, including: When segmenting Chinese proper nouns, due to their complex structure, standard word segmentation often mis-segments entities, resulting in incomplete entity recognition. The recognition ability of long entities is weak, especially in scenarios lacking context or prior knowledge; The external models and dictionaries they carry have high maintenance costs and are difficult to adapt to domain texts with frequent semantic changes or low-resource application scenarios.

[0004] The present invention provides a named entity extraction method based on term merging to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a named entity extraction method based on term merging, which automatically identifies named entities through structural merging based on term combinations and frequency statistics after word segmentation, without the need for any external labeled data or pre-trained models.

[0006] In order to achieve the above objectives, the basic solution of the present invention provides a method for named entity extraction based on term merging, comprising the following steps: Step A1: Processing the original text to obtain a global term set including a plurality of term lists divided in sentence order, wherein the term lists include a plurality of word segments divided in character string order; Step A2: Create a word segmentation position index based on the initial string position; Step A3: combine the token at the current token position index and the token at the next string position to form a token group, check whether the string length of the token group meets the minimum word length limit, if yes, jump to step A6, otherwise, check the token merging condition of the token group, if yes, go to step A4, otherwise, jump to step A6; Step A4: merge the token group into a long entity and add it to the final entity set, replace the token at the string position corresponding to the original token position index with the token group as a new token, and delete the token at the next string position; Step A5: when the length of the current token position index is less than the length of the current token list minus one, repeat steps A3 to A5 based on the current token position index, otherwise go to step A7; Step A6: move one string position backward to establish a new token position index, when the length of the new token position index is greater than the length of the current token list minus one, go to step A7, otherwise, go back to step A3 and loop through steps A3 to A6; Step A7: repeat steps A3 to A7 until all token lists are traversed, output the final entity set.

[0007] Further, the tokenization processing of the original text includes sentence-level segmentation of the original text according to punctuation marks, and processing of each segmented sentence using a general Chinese tokenization method, finally obtaining a global token set composed of all tokens in string order.

[0008] Further, the position of the token position index corresponding to the initial string position is 0.

[0009] Further, the default value of the minimum word length limit is 2.

[0010] Further, the token merging condition is: Count the global occurrence frequency of the token group in the global token set, and set a minimum frequency threshold as the merging condition. When the global occurrence frequency is greater than or equal to the minimum frequency threshold, it means that the token merging condition is met, otherwise it means that the token merging condition is not met.

[0011] Further, the specific counting method of the global occurrence frequency is as follows: Traverse each token list in the global token set one by one, and extract all continuous token pairs in the process using a sliding window method. Compare the extracted continuous token pairs with the token group one by one, if a match is found, increment the frequency counter by one, until all token lists are traversed. The count value of the frequency counter is the global occurrence frequency of the token group under the current token set.

[0012] Further, the window size of the sliding window is 2 and the step size is 1.

[0013] Further, in step A3, the minimum frequency threshold value is 5 by default.

[0014] The principle and effect of the basic scheme are that: Compared with the prior art, the present application realizes the extraction of named entities by automatically mining and merging continuous and high-frequency word combination from the original text in sequence without relying on external dictionaries or training models through the following steps: sequential combination judgment, full-text frequency statistics, and word merging. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 A flowchart of a named entity extraction method based on word merging according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the drawings and preferred embodiments.

[0018] A named entity extraction method based on word merging, as shown in Figure 1 The method comprises the following steps: Step A1: sequentially performing sentence-level segmentation and word segmentation on the original text to be subjected to named entity extraction to obtain a global word set W_all composed of all word segmentation results, the global word set W_all comprising a plurality of word lists, the word lists being in the same order as the sentence order, and each word list being composed of the word segmentation results of the corresponding sentence after word segmentation arranged in the order of character strings.

[0019] The global word set W_all is specifically represented as follows: ], represents the word list composed of all word segmentation results of the kth sentence, , [ ], represents the word segmentation result of the kth sentence in the order of character strings in the corresponding word list, m .

[0020] Specifically, word segmentation processing of the original text involves segmenting the original text at the sentence level according to punctuation marks, and then applying a common Chinese word segmentation method to each segmented sentence, ultimately obtaining a global word set consisting of all segmented words. Punctuation marks include, but are not limited to, periods, question marks, and exclamation marks.

[0021] For example, when the original text obtained is: "The article was published in 1987 and the author is very old"; The segmentation results after word segmentation are: “The article was published in 1987 and the author is very old”; The corresponding global term set W_all is expressed as follows: [ [Article, published, in, 1987, year], [Editor, age, very high] ].

[0022] Step A2: Create a word segmentation position index j based on the initial string position; Step A3: Starting from the initial word segmentation position index j=0, The word segmentation at the corresponding string position in and the word segmentation of the next string position Merge and combine to form word groups[ ]; Establish the minimum word length limit min_len after merging, and combine it with the word group [ ] string length, when the word group [ ] is less than the minimum word length min_len, jump to step A6. Otherwise, count the global occurrence frequency f of the word group in the global word set W_all, and establish a minimum frequency threshold freq as the word merging condition. When the global occurrence frequency f is greater than or equal to the minimum frequency threshold freq, it means that the word merging condition is met and go to step A4. Otherwise, jump to step A6.

[0023] In this embodiment, the established minimum word length limit is expressed as min_len, and min_len=Term_min is set. The type of Term_min is int, and the default value of the minimum word length limit min_len is 2.

[0024] In this embodiment, for word groups greater than the minimum word length limit min_len, ], the specific statistical method of its global occurrence frequency f in the global term set W_all is as follows: Traverse each term list in the global term set W_all one by one In this process, all continuous word pairs are extracted by sliding window, with the window size set to 2 and the step size set to 1. The extracted continuous word pairs are sequentially compared with the word group [ ] is compared. If the match is successful, the frequency counter is increased by one until all the word lists are traversed. The count value of the obtained frequency counter is the global occurrence frequency f of the word group in the current word set.

[0025] In this embodiment, the established minimum frequency threshold is represented by freq, the type is int, the default value is 5, and min_freq=freq.

[0026] Step A4: The word groups that meet the merging conditions [ ] are merged into a long entity and added to the final entity set entity_set, and in the list of terms The middle participle group[ ] into a new term , and by this new term Replace the original participle in the belonging term list , including Strings merged into The position of the string, while deleting String.

[0027] In this example, consecutive terms are also synchronously merged in all term sublists in the global term set W_all: ], when it appears in other places If there is a combination of , it will be merged. When traversing to this position, there is no need to recalculate.

[0028] Step A5: Count the length of the current word list, record it as n, if j <n-1,则在保持起始位置j不变的情况下,重复步骤A3至步骤A5,重新判断当前位置新的分词组[ ]; Otherwise, it indicates that the current list of terms has been judged, and the process goes to step A7; Step A6: Move back one string position to create a new word segmentation position index, let , and count the length of the current word list, recorded as n, if >n-1, indicating that the current word list is judged and the process goes to step A7. <n-1,回到步骤A3并循环遍历步骤A3至步骤A6,直至当前词项列表中所有可以合并的、连续的分词合并完毕; Step A7: according to the word list arrangement order in the global word set W_all, the (k+1)th word list is selected in turn to repeat steps A3 to A7, until all word lists are traversed, the judgment of the sentence of all word lists is completed, and the final entity set entity_set is output. The type of the final entity set entity_set is list.

[0029] The naming entity recognition method based on word merging provided by the application realizes the extraction of the naming entity by automatically mining and merging the continuous and high-frequency word combination from the original text through the series of steps of sequential combination judgment, full-text frequency statistics and word merging without relying on external dictionaries or training models.

[0030] Compared with the existing naming entity recognition method, the application has significant advantages in many aspects: (1) No external dependence, strong adaptability: without relying on any external dictionary, pre-training model or artificial annotation corpus, it is suitable for open domain, low resource or real-time text processing scene; (2) Low consumption of computing resources, suitable for lightweight application: without model loading or training, it can run efficiently on terminal or edge device with limited resources.

[0031] The above is only the preferred embodiment of the application, and does not limit the application in any form. Although the application has been disclosed as above, it is not intended to limit the application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the technical solution of the application, and any indirect modification, equivalent change and modification of the above embodiment according to the technical essence of the application still belong to the scope of the technical solution of the application.

Claims

1. A named entity extraction method based on term merging, characterized in that: The steps include: Step A1: Processing the original text to obtain a global term set including a plurality of term lists divided in sentence order, wherein the term lists include a plurality of word segments divided in character string order; Step A2: Create a word segmentation position index based on the initial string position; Step A3: The word segmentation indexed by the current word segmentation position and the word segmentation indexed by the next string position are combined into a word group. The string length of the word segmentation group is checked to see if it meets the minimum word length limit. If so, the process jumps to step A6. Otherwise, the word segmentation group is checked for the term merging condition. If so, the process jumps to step A4. Otherwise, the process jumps to step A6. Step A4: Merge the word group into a long entity and add it to the final entity set. Use the word group as a new term to replace the word at the string position corresponding to the original word position index, and delete the word at the next string position. Step A5: When the length of the current segmentation position index is less than the length of the current term list minus one, repeat steps A3 to A5 based on the current segmentation position index; otherwise, proceed to step A7; Step A6: Move the string position backward by one to establish a new word segmentation position index. If the length of the new word segmentation position index is greater than the length of the current word list minus one, proceed to step A7. Otherwise, return to step A3 and loop through steps A3 to A6. Step A7: Repeat steps A3 to A7 until all term lists are traversed and the final entity set is output.

2. A method for named entity extraction based on term merging according to claim 1, characterized in that: The word segmentation processing of the original text includes sentence-level segmentation of the original text according to punctuation marks, and processing each sentence after segmentation using a common Chinese word segmentation method, and finally obtaining a global word set consisting of all segmented words in string order.

3. The method for named entity extraction based on term merging according to claim 1, characterized in that: The word position index corresponding to the initial string position is 0.

4. The method for named entity extraction based on term merging according to claim 1, characterized in that: The default value of the minimum word length limit is 2.

5. The method for named entity extraction based on term merging according to claim 1 or 4, characterized in that: The term merging conditions are: The global occurrence frequency of the word groups in the global word set is counted, and a minimum frequency threshold is established as a merging condition. When the global occurrence frequency is greater than or equal to the minimum frequency threshold, it means that the word merging condition is met, otherwise it means that the word merging condition is not met.

6. The method for named entity extraction based on term merging according to claim 5, characterized in that: The specific statistical method for global occurrence frequency is as follows: Traverse each term list in the global term set one by one, and use a sliding window method to extract all consecutive term pairs. The extracted consecutive term pairs are compared with the word groups in turn. If the match is successful, the frequency counter is increased by one until all term lists are traversed. The count value of the obtained frequency counter is the global occurrence frequency of the word group in the current term set.

7. The method for named entity extraction based on term merging according to claim 6, characterized in that: The window size of the sliding window is 2 and the step size is 1.

8. The method for named entity extraction based on term merging according to claim 6, characterized in that: In step A3, the default value of the minimum frequency threshold is 5.

Citation Information

Patent Citations

  • Keyword extraction method integrating theme information and bidirectional LSTM

    CN109933804A

  • Entity recognition method and device based on semantic analysis, equipment and storage medium

    CN111368547A

  • Retrieval method, retrieval device, computer readable medium and electronic equipment

    CN112445959A

  • Named entity recognition method, device and equipment and computer readable storage medium

    CN113011186A

  • HAZOP named entity recognition and entity relationship extraction method

    CN118761406A