Chinese text nested named entity recognition method based on soft hint
Through the PSNER model, the complexity problem of nested named entity recognition in the financial field is solved by using soft prompts and scale transformation self-attention mechanisms, and the accurate identification of multi-level nested structures and entity boundaries is achieved, which improves the recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202510412091.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
The existing nested named entity recognition methods are difficult to effectively deal with complex multi-level nested structures and difficult-to-define entity boundaries in the financial field, resulting in poor recognition results.
The nested named entity recognition model PSNER in the Chinese financial field based on soft prompts is adopted. By introducing soft prompts and scale-transformed self-attention mechanisms, the context position information and spatial characteristics of the span are fully utilized, and combined with the BERT encoder, BiLSTM network and the contrast learning loss function, the semantic dependency coding capability between entities is enhanced.
It improves the accuracy and efficiency of nested entity recognition, can effectively handle multi-level nested structures and complex entity boundaries in the financial field, and enhances the model's semantic dependency coding ability between entities.
Smart Images

Figure CN120354852A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a method for identifying nested named entities in the Chinese financial field based on soft prompts, which is particularly suitable for processing financial text entity recognition tasks with multi-level nested structures. Background Art
[0002] Nested named entity recognition aims to accurately identify multi-level and specifically meaningful entities, such as person names, place names, organization names, etc., from complex unstructured text data. This in-depth entity recognition can more accurately understand the semantic information in the text and is crucial for many applications in the field of natural language processing (Natural Language Processing, NLP), such as information extraction, relation extraction, and question answering systems. However, traditional named entity recognition (Named Entity Recognition, NER) methods mainly focus on the recognition of entities with a flat structure and often have difficulty effectively processing complex nested structures. The financial field, as an information-intensive industry, contains rich valuable information in its text data, but at the same time, it also contains a large number of nested entities. Figure 1 Fig. shows a four-layer nesting example containing two entity types: place names and company names. This complex nested structure requires the model to accurately distinguish entities at different levels, greatly increasing the difficulty of entity recognition. In addition, nested entities can effectively represent the semantics of named entities and are widely used in downstream tasks. Therefore, the nested named entity recognition task has become one of the current research hotspots.
[0003] For the nested named entity recognition task, existing methods can mainly be divided into methods based on sequence labeling, rules, hypergraphs, and spans. These methods have made significant progress in the nested named entity recognition task, but there are also certain limitations. Methods based on sequence labeling are difficult to directly represent nested entities because a token may belong to multiple named entities simultaneously, resulting in poor recognition performance; methods based on rules and hypergraphs rely on a large number of manual features, which are not only inefficient but also difficult to handle complex and variable text data. Among them, the span-based nested named entity recognition method is relatively popular. This method can flexibly process entities of different lengths and has good adaptability to nested entities. However, they often assume that spans are independent and ignore the context information and spatial features of spans. Context information is crucial for understanding the semantics of entities and determining their boundaries, while spatial features help capture the hierarchical relationships and nested structures between entities. The lack of this information undermines the semantic dependency encoding ability between spans, leading to possible misjudgments or omissions when the model recognizes nested entities. In addition, most of these methods are applicable to the general domain, but face many challenges in financial domain applications: 1) The nested entity structure in the financial domain is relatively complex, and there are often multi-level nested relationships between entities. For example, in the financial domain dataset CFNE we collected, nested entities account for about 34.2% of all entities, and entities with more than three levels of nesting account for about 6.5% of all entities. This requires the recognition method to be able to understand the hierarchical relationships and nested structures between entities. 2) It is more difficult to define the boundaries of entities in the financial domain. Due to the high professionalism and complexity of financial domain entities, such as the mixed use of company full names and abbreviations, and professional terms in the financial domain, it is more difficult to define the boundaries of nested entities in the financial domain. Summary of the Invention
[0004] The present invention provides a soft prompt-based Chinese financial domain nested named entity recognition model PSNER to solve the problems that the nested entity structure in the financial domain is relatively complex, there are often multi-level nested relationships between entities, and it is more difficult to define the boundaries of entities in the financial domain.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] Design a soft prompt-based Chinese financial domain nested named entity recognition model (PSNER), including the following steps:
[0007] S1: Screen out the outermost entities contained in the input Chinese financial domain text, and perform word segmentation to obtain a candidate nested span set;
[0008] S2: For each sentence, insert soft prompts before and after the candidate spans to generate an encapsulated text sequence;
[0009] S3: Use the pre-trained language model BERT to encode the encapsulated text sequence to obtain context embedding representations
[0010] S4: Combine word embeddings character embeddings and context embeddings to generate the final token representation h through a bidirectional long short-term memory network i ;
[0011] S5: Perform max pooling operation on each candidate span to generate a feature representation h p , and combine the left and right boundary features h l-1 、h r+1 to generate a span representation span i ;
[0012] S6: Introduce a contrastive learning loss function to shorten the semantic distance between positive sample pairs;
[0013] S7: Through the scale-transformed self-attention mechanism, fuse the width and center position features of the span to generate an enhanced span representation;
[0014] S8: Classify the enhanced span representation and output the nested entity type and boundaries.
[0015] Furthermore, in step S1, we first filter out the outermost entities contained in the text and tokenize the text, denoted as the M module. According to the tokenization results and the outermost entities, a candidate nested span set is obtained, denoted as S=(s1, s2,..., s w ).
[0016] Furthermore, in step S2, for each sentence, soft prompts are inserted before and after the candidate span respectively, and it is encapsulated as x p ={x part1 , [p1], s i , [p2], x part2}, where [p1] represents the soft prompt. For the sentence "On the innovative services of Bank of China", it is encapsulated into three sequences "On [p1]China [p2]Bank's innovative services", "In [p1]Bank [p2] of China's innovative services", and "On [p1]Bank of China [p2]'s innovative services".
[0017] Furthermore, in step S3, this paper uses BERT as a text encoder to obtain context embeddings
[0018]
[0019] Furthermore, in step S4, the pre-released resource file is used to obtain word embeddings, and the context embeddings are concatenated with the word embeddings and the character embeddings . Then they are input into a BiLSTM layer, and the final representation of the i-th token is:
[0020]
[0021] h i = BiLSTM([f1, f2,..., f m )
[0022] where represents the concatenation operation.
[0023] Furthermore, in step S5, for the representation of the span, since the soft prompt contains context or semantic information about the entity and can be adaptively adjusted according to the training data, therefore, this paper also integrates this information h l-1 , h r+1 into the representation of the span:
[0024] h p = MaxPooling(h l ,..., h r )
[0025]
[0026] where h p represents the feature representation after max pooling.
[0027] Furthermore, in step S6, the construction of the soft prompt often causes potential interference to nested entities, resulting in the disconnection of the connection between entities that should be closely related. To alleviate this problem, contrastive learning is introduced to handle it. For the same sentence text, the encapsulated sequence x pi , is used as the positive sample pair (x pi , ). Since in the nested named entity recognition task, excessive attention to negative samples may cause the model to over-focus on irrelevant or opposite information and ignore the key connection between positive sample pairs, and this paper is more concerned about how to strengthen the connection between positive sample pairs, so negative sample pairs are not constructed. This paper uses cosine similarity to construct a contrastive loss function, which aims to maximize the similarity between positive sample pairs:
[0028]
[0029] where x piis the CLS representation obtained by BERT, represents the cosine similarity, and τ is the temperature hyperparameter.
[0030] Furthermore, in step S7, a scaled transformation self-attention mechanism is adopted to enhance the representation of the span. For span s i =(l, r), define the span width w = r - l + 1, and define the span center position Integrate these two geometric information into the span representation to generate two correlation matrices to reflect the connection between each span. Calculate the ratio of the overlap degree between two spans to the smaller width among them. If two spans do not overlap, the correlation is 0; if they completely overlap, the correlation is 1. Add 1 to the Euclidean distance to avoid division by zero, and take its reciprocal to obtain the correlation measure between the center positions. The larger the value, the closer the center positions of the two spans are, and the higher the correlation; conversely, the lower the correlation:
[0031]
[0032] where c i represents the center position of the i-th span, and w i represents the width of the i-th span. d(c i , c j ) is the Euclidean distance between the center positions of the i-th span and the j-th span.
[0033] Furthermore, connect the width correlation matrix and the center position correlation matrix to obtain the scaled transformation matrix M st :
[0034]
[0035] Furthermore, calculate the enhanced span representation:
[0036]
[0037] where w ij is a weight matrix.
[0038] Finally, use an MLP to classify the span:
[0039]
[0040] Adopt weighted cross-entropy loss:
[0041]
[0042] where w iis the weight of the i-th span; level(i) is the nesting level of the i-th span; freq(i) is the frequency of the i-th span appearing in the training data; α is a hyperparameter used to balance the influence of the nesting level and entity frequency on the weight.
[0043] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0044] 1. The present invention fully utilizes the context position information of spans by introducing soft prompts and integrates the soft prompts into the span representation, making full use of the potential connection information between nested entities. The introduction of soft prompts enables the model to capture the context position information of spans and integrate it into the span representation, which not only helps the model better understand the semantics of entities but also makes full use of the potential connection information between nested entities to enhance the semantic dependency encoding ability between entities. In addition, BERT is used as a text encoder to obtain context embeddings, and the final token representation is generated through a bidirectional long short-term memory network by combining word embeddings and character embeddings, realizing the organic integration of global semantics, local context, and morphological features, providing an efficient and reliable technical foundation for the structured processing of financial texts. And on this basis, contrastive learning is introduced, which helps to alleviate the potential interference caused by soft prompt construction to nested entities.
[0045] 2. The present invention introduces a scale-transformed self-attention mechanism to integrate spatial features into the representation of candidate spans, enhancing the model's ability to capture the interaction relationship between spans. The introduction of the scale transformation module enables the model to capture the interaction relationship between spans and integrate spatial features to enhance the span representation, so as to capture the hierarchical relationship and nested structure between entities.
[0046] 3. The present invention completes the recognition of nested entities in Chinese financial text under the condition that the nested entity structure in the financial field is relatively complex and the entity boundaries in the financial field are more difficult to define, which is of great significance for tasks such as information extraction and relationship mining. Description of the Drawings
[0047] Figure 1 is an example of nested entities Figure 1 .
[0048] Figure 2 is the PSNER model architecture Figure 2 .
[0049] Figure 3 is the statistical data table 1 of three nested named entity recognition datasets.
[0050] Figure 4 is the result table 2 of the PSNER model and the baseline model on the financial dataset CFNE.
[0051] Figure 5Table 3 shows the results of the present PSNER model and the baseline model on the CMeEE and CNERTA datasets.
[0052] Figure 6 Table 4 shows the results of the ablation experiment.
[0053] Figure 7 The impact of different numbers of training samples on the model Figure 3 。
[0054] Figure 8 Table 5 shows the case study.
[0055] The following describes the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments. However, the following embodiments are only used to illustrate the present invention in detail and do not limit the scope of the present invention in any way.
[0056] Figure 1 Figure 1 shows a four-layer nesting example containing two entity types: location and company name. This complex nesting structure requires the model to accurately distinguish entities at different levels, greatly increasing the difficulty of entity recognition.
[0057] Embodiment 1: A Chinese text nested named entity recognition model (PSNER) based on soft prompts. Refer to Figure 2 , which consists of two parts: the candidate entity extraction part and the span classification part.
[0058] 1) Candidate entity extraction part: Refer to the left half of Figure 2 . The nested NER is formalized into candidate entity extraction and span classification to obtain candidate spans and tokenization results. Then, soft prompts are inserted before and after the candidate spans to make full use of the context position information, and contrastive learning is used to shorten the distance between sentence representations. First, the outermost entities contained in the text are screened out, and the text is tokenized. According to the tokenization results and the outermost entities, a candidate nested span set is obtained. For each sentence, soft prompts are inserted before and after the candidate spans, and then BERT is used as the text encoder to obtain context embeddings. Combining word embeddings, character embeddings, and context embeddings, the final token representation is generated through a bidirectional long short-term memory network. A max-pooling operation is performed on each candidate span, and span representations are generated by combining the left and right boundary features of the span. Finally, a contrastive learning loss function is introduced to shorten the semantic distance between positive sample pairs.
[0059] 2) Span classification part: Refer to Figure 2In the right half, the scale-transformed self-attention mechanism is used to enhance the span representation and complete span classification. First, two geometric information, namely the span width and the span center position, are incorporated into the span representation to generate two correlation matrices to reflect the connections between each span; calculate the ratio of the overlap degree between two spans to the smaller width of them; finally, the width-related matrix and the center-position-related matrix are concatenated to obtain the scale-transformed matrix, and the enhanced span representation is calculated.
[0060] The present invention is divided into two parts: extracting candidate entities and span classification. In the candidate entity extraction module, soft prompts are introduced to make full use of the contextual position information of the spans, and the soft prompts are incorporated into the span representation to make full use of the potential connection information between nested entities. In the span classification module, the scale-transformed self-attention mechanism is adopted to integrate spatial features into the representation of candidate spans, enhancing the model's ability to handle the interaction relationships between spans.
[0061] The present invention conducts experiments on three nested named entity recognition datasets: CMeEE, CNERTA, and the nested dataset CFNE re-annotated on a dataset, Figure 3 showing the statistics of these datasets. The CFNE dataset is a nested named entity recognition dataset re-annotated from a certain financial dataset, including six entity types: PERSON_NAME, LOCATION, TIME, ORG_NAME, COMPANY_NAME, and PRODUCT_NAME. Among them, nested entities account for about 34.2% of all entities. The CMeEE dataset is a nested named entity recognition dataset for Chinese medical texts, coming from the well-known Chinese medical NLP evaluation benchmark CBLUE. This dataset collects a large amount of medical text data and ensures the accuracy of the data through a series of strict data cleaning and annotation processes. And before annotation, the text has been automatically segmented to ensure that all medical entities have been correctly segmented. The dataset divides the named entities in medical texts into nine categories, including: disease (dis), clinical manifestation (sym), drug (dru), medical device (equ), medical procedure (pro), body (bod), medical test item (ite), microorganism (mic), department (dep). Among them, nested entities account for about 10.7% of all entities. The CNERTA dataset is a large-scale manually annotated Chinese multi-modal named entity recognition dataset, covering 42,987 annotated sentences and 71 hours of speech data. In this paper, we only use the text content of this dataset, which contains three entity types: person name, location, and organization. Among them, nested entities account for about 28.2% of all entities.
[0062] The present invention uses BERT as a text encoder to obtain context embeddings and Wiki to obtain word embeddings. For the medical dataset CMeEE, BERT is replaced by BioBERT. The dimension of the hidden layer is 1024, the batch size is set to 64, the number of epochs is set to 50, and the learning rate is set to 10 -3 , and the hyperparameter τ is set to 0.08.
[0063] In the present invention, BERT is used as a text encoder to obtain context embeddings. The context embeddings are concatenated with word embeddings and character embeddings, and then they are input into a BiLSTM layer. The final representation of the i-th token is:
[0064]
[0065] h i = BiLSTM([f1, f2,..., f m )
[0066] where represents the context embeddings, which are obtained by using BERT as a text encoder; represents the word embeddings, represents the character embeddings, represents the concatenation operation.
[0067] Regarding the representation of spans in the present invention, since the soft prompts contain context or semantic information about entities and can be adaptively adjusted according to the training data, the present invention also integrates this information h l-1 , h r+1 into the representation of spans:
[0068] h p = MaxPooling(h l ,..., h r )
[0069]
[0070] where h p represents the feature representation after max pooling.
[0071] To alleviate the potential interference of the construction of soft prompts on nested entities, the present invention introduces contrastive learning and uses cosine similarity to construct a contrastive loss function, which aims to maximize the similarity between positive sample pairs:
[0072]
[0073] where x piis the CLS representation obtained by BERT, represents the cosine similarity, and τ is the temperature hyperparameter.
[0074] The present invention uses a scale transformation self-attention mechanism to enhance the representation of spans, incorporates two geometric information, namely span width and span center position, into the span representation, and generates two correlation matrices to reflect the connections between various spans. Calculate the ratio of the overlap degree between two spans to the smaller width among them. If the two spans do not overlap, the correlation is 0; if they completely overlap, the correlation is 1. Add 1 to the Euclidean distance to avoid division by zero, and take its reciprocal to obtain the correlation measure between the center positions. The larger the value, the closer the center positions of the two spans are, and the higher the correlation; conversely, the lower the correlation:
[0075]
[0076]
[0077] where c i represents the center position of the i-th span, and w i represents the width of the i-th span. d(c i , c j ) is the Euclidean distance between the center positions of the i-th span and the j-th span.
[0078] The present invention connects the width correlation matrix and the center position correlation matrix to obtain the scale transformation matrix M st , and calculates the enhanced span representation:
[0079]
[0080] where w ij is a weight matrix.
[0081] Finally, use an MLP to classify the spans and adopt weighted cross-entropy loss:
[0082]
[0083] where w i is the weight of the i-th span; level(i) is the nesting level of the i-th span; freq(i) is the frequency of the i-th span appearing in the training data; α is a hyperparameter used to balance the influence of nesting level and entity frequency on the weight.
[0084] Through Figure 4 and Figure 5 comparative experiments, it can be seen that the present invention has better performance in the nested entity recognition task in the financial field. ThroughFigure 6 The ablation experiment shows the importance of the soft prompt module and the scale transformation module in the present invention.
[0085] Through Figure 7 , the present invention further explores the influence of the size of the training samples on the model performance. It can be seen from the figure that regardless of the change in the amount of training data, the PSNER model always maintains its advantage over the baseline model; through Figure 8 , the present invention verifies the effectiveness of the PSNER model in learning text semantics and structure in specific cases, and can better capture the semantic dependencies in the text, so as to achieve effective recognition of nested entities.
[0086] The present invention has been described in detail above in conjunction with the accompanying drawings and embodiments. However, those skilled in the art can understand that without departing from the purpose of the present invention, various specific parameters in the above embodiments can be changed to form multiple specific embodiments, which are all within the common change range of the present invention and will not be elaborated herein one by one.
Claims
1. A method for identifying nested named entities in Chinese text based on soft prompts, characterized in that, It includes the following steps: S1: Screen out the outermost entities contained in the input Chinese financial domain text, and perform word segmentation to obtain a candidate nested span set; S2: For each sentence, insert soft prompts before and after the candidate spans to generate an encapsulated text sequence; S3: Use the pre-trained language model BERT to encode the encapsulated text sequence and obtain the context embedding representation S4: Combine word embeddings Character embeddings and context embeddings to generate the final token representation h through a bidirectional long short-term memory network i ; S5: Perform max pooling operation on each candidate span to generate a feature representation h p , combine the left and right boundary features h l-1 、h r+1 to generate a span representation span i ; S6: Introduce a contrastive learning loss function to shorten the semantic distance between positive sample pairs; S7: Through a scale-transformed self-attention mechanism, fuse the width and center position features of the spans to generate an enhanced span representation; S8: Classify the enhanced span representation and output the nested entity type and boundaries.
2. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein In step S1, first, the outermost entities included in the text are screened out, and the text is tokenized, denoted as module M. According to the tokenization results and the outermost entities, a candidate nested span set is obtained, denoted as S=(s1, s2,..., s w ).
3. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein In step S2, for each sentence, soft prompts are inserted before and after the candidate span, and it is encapsulated as x p = {x part1 , [p1], s i , [p2], x part2}, where [p1] represents the soft prompt. For the sentence "On the innovative services of Bank of China", it is encapsulated into three sequences "On [p1]China [p2]Bank of the innovative services", "On China [p1]Bank [p2] of the innovative services", and "On [p1]Bank of China [p2] of the innovative services".
4. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein In step S3, this paper uses BERT as a text encoder to obtain context embeddings 5. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein, In step S4, a pre-released resource file is used to obtain word embeddings, and the context embeddings are concatenated with the word embeddings and character embeddings, and then they are input into a BiLSTM layer.
6. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein In steps S5 and S6, the soft prompt is an adaptively adjusted token, which is used to capture the context position information of the spans and reduce its interference with the original semantics through contrastive learning.
7. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein In step S7, the scale-transformed self-attention mechanism is implemented through the following steps: S1: Calculate the width correlation matrix of two spans, which reflects the overlap degree between the spans; S2: Calculate the center position correlation matrix of the spans, which reflects the spatial distance between the spans; S3: Concatenate the width and center position correlation matrices to generate a scale transformation matrix; S4: Based on the scale transformation matrix, perform weighted fusion on the span representations.
8. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein, The contrastive learning loss function is realized by maximizing the cosine similarity of positive sample pairs, and no negative sample pairs are constructed.
9. The method for identifying nested named entities in Chinese text based on soft prompts according to claim 1, wherein, During the training process, jointly optimize the contrastive loss and the entity classification loss, and balance the influence of nested levels and entity frequencies through weighted cross-entropy loss.