Hierarchical pointer nested entity recognition model and method based on cross attention

By adopting a model of cross attention and Drop-up feature fusion in nested named entity recognition, the problems of blurred entity boundaries and low feature adhesion are solved, and clearer entity boundaries and higher recognition performance are achieved.

CN120218070AInactive Publication Date: 2025-06-27XICHANG NAT PRESCHOOL TEACHERS COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510300181.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing nested named entity recognition methods have problems such as blurred entity boundaries, low feature adhesion, and failure to effectively utilize multi-statement sequences.

Method used

A hierarchical pointer nested entity recognition model based on cross attention is adopted, and the feature representation and entity boundary representation are strengthened through cross-position coding and Drop-up feature fusion method, and a Pro-Clash method is designed to solve the conflict problem between entities.

Benefits of technology

It improves the clarity of entity boundaries, the robustness and generalization of the model, and can extract nested entities more effectively, with better performance than current mainstream models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218070A_ABST
    Figure CN120218070A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, discloses a hierarchical pointer nested entity recognition model and method based on cross attention, and aims to solve the problems that in a nested named entity recognition method, the feature adhesion degree is low, the boundary is fuzzy, and a multi-statement sequence cannot be effectively utilized. The invention provides a nested entity recognition method fusing cross attention and a feature Drop-up. Firstly, position information is added for characters through designed cross position codes in a word embedding layer, and boundary representation is enhanced; secondly, mixing the features by using the extracted Drop-up so as to enhance feature interaction; designing a multi-layer entity identification framework based on pointer network identification, and judging entity head and tail indexes; finally, a Pro-Cash method is designed to solve the problem of conflicts among entities, and generalization and recognition performance are comprehensively improved. Moreover, the method provided by the invention is generally superior to a current mainstream method, and has a good recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and more specifically, to a hierarchical pointer nested entity recognition model and method based on cross-attention. Background Art

[0002] Named entity recognition is a basic task in natural language processing (NLP), aiming to identify proper nouns such as person names, place names, and organization names from text sequences, extract key information, and provide information support for downstream tasks such as relational syntactic analysis, text classification, machine translation, and knowledge graphs. Therefore, it has important research significance in NLP tasks.

[0003] Due to the semantic diversity and complexity in natural text sequences, there are often many proper nouns with nested structures. Such nested entities can better assist tasks such as intelligent question answering and event extraction. Therefore, nested named entity recognition is not only a major component of named entity recognition research but also a key and difficult point to be solved.

[0004] For the research on nested named entity recognition, early methods used strategies such as hypergraphs and multi-layer sequence annotation to start from flat entities and extract nested entities successively, but they had high complexity and lost fine-grained information. In contrast, the span-based method is easy to implement and has fine classification, becoming the current mainstream research method.

[0005] However, existing methods still face problems such as fuzzy entity boundaries, insensitivity to distance information between entities, and conflicts between multi-layer entities. Summary of the Invention

[0006] In order to overcome the above-mentioned defects existing in the prior art, the present invention provides a hierarchical pointer nested entity recognition model and method based on cross-attention. Aiming at the problems of low feature adhesion, fuzzy boundaries, and failure to effectively utilize multi-sentence sequences in nested named entity recognition methods, a nested entity recognition method that combines cross-attention and feature Drop-up is proposed. First, at the word embedding layer, position information is added to characters through the designed cross-position encoding to strengthen the boundary representation; secondly, the proposed Drop-up is used to confuse the features to enhance feature interaction; then, based on the pointer network, a multi-layer entity recognition framework is designed to judge the start and end indexes of entities; finally, a Pro-Clash method is designed to solve the problem of entity conflicts, comprehensively improving the generalization and recognition performance.

[0007] The above technical objectives of the present invention are achieved through the following technical solutions: A hierarchical pointer nested entity recognition model and method based on cross-attention, including an embedding layer, a feature interaction layer, a decoding layer, and a prediction layer;

[0008] The embedding layer uses RoBERTa as the character encoding layer;

[0009] The feature interaction layer injects position information and confuses features for the output of the embedding layer through cross-position encoding and Drop-up feature fusion method to strengthen the feature representation;

[0010] The decoding layer builds an entity category layer based on the pointer network and optimizes the parameters using a multi-label classification loss function;

[0011] The prediction layer filters out the position labels less than 0 for the logits obtained by the decoding layer, and then traverses to collect the complete entity set.

[0012] The hierarchical pointer nested entity recognition method based on cross-attention is characterized by including the following steps:

[0013] S1. Inject position information into the character vector through cross-position encoding to obtain a relative position vector;

[0014] S2. Confuse the relative position vector and the initial character embedding through Drop-up feature confusion to strengthen the feature representation;

[0015] S3. Improve the pointer network structure, use the start and end classification matrices to mark the candidate entity set, and introduce a multi-label cross-loss entropy function to quickly train and reduce error propagation;

[0016] S4. Adopt the Pro-Clash method to filter the entities with interactions in the candidate intervals to further solve the entity conflict problem.

[0017] Furthermore, the character vector is obtained by encoding the characters and their corresponding positions in the text sequence into RoBERTa;

[0018] The expression of the text sequence is as follows:

[0019] S = {x1, x2, x3... x n} (1)

[0020] In the formula, x n represents the character in the text sequence, and n represents the corresponding position;

[0021] The formula of RoBERTa is as follows:

[0022] RoBERTa_out = RoBERTa(S) (2)

[0023] Furthermore, the cross-position encoding specifically includes the following steps:

[0024] ① The RoPE model is used to achieve the goal of including distance information while taking the inner product of two vectors. The formula is as follows:

[0025] <f q (q m ,m),f k (k n ,n)>=g(q m ,k n ,m - n) (3)

[0026] In the formula, m - n represents the distance information, and q m 、k n represent the query vector and the key vector at positions m and n;

[0027] ② Add the distance information of m - n to the inner product between positions m and n. The formula is as follows:

[0028] f q (q m ,m)=q m e imθ =(W q x m )e imθ

[0029] f k (k n ,n)=k n e inθ =(W k x n )e inθ (4)

[0030] g(q m ,k n ,m - n)=Re[q m k n * e i(m-n)θ =Re[(W q x m )(W k x n ) * e i(m-n)θ

[0031] In the formula, e imθ 、e inθ are complex exponential functions represented by Euler's formula, which are related to trigonometric functions and expanded into rotation matrices, where i is the imaginary unit and e is the base of the natural logarithm; k n * is the conjugate complex number of k n , and Re represents taking the real part;

[0032] ​Among them, the expression of the value range of θ is as follows:

[0033] θ = [θ0,..., θ d / 2-1 (5)

[0034] In the formula, d is the dimension length of the query vector, which is used for calculation by substituting into the sine and cosine functions;

[0035] ③ Perform rotational position encoding on the embedding vector, and at the same time perform the reverse operation on the query vector. The formula is as follows:

[0036] q m = contrary(RoBERTa_out)·W Q

[0037] k n = RoBERTa_out·W K (6)

[0038] V = RoBERTa_out·W V

[0039] In the formula, RoBERTa_out is the embedding vector, and contrary is the reverse operation to reverse-encode the query vector;

[0040] ④ Use the position encoding in RoPE to multiply the query and key vectors by the rotation matrix for calculation. The formula is as follows:

[0041]

[0042] ⑤ Perform dot product attention calculation on the query and key vectors after position encoding and the value vector to obtain the final position feature CrossPE_out. The formula is as follows:

[0043]

[0044] Furthermore, in step S2, the formula for Drop-up feature confusion is as follows:

[0045]

[0046] In the formula, respectively represent the embedding vector RoBERTa_out and the position feature CrossPE_out. Dropout = 1 means that the encoding vector t1 of character i is set to 0 using Dropout, otherwise the t2 vector at the corresponding position is output.

[0047] Further, in step S3, the method for improving the pointer network structure is specifically as follows: On the basis of the pointer network, an entity category layer is built, and a multi-label classification loss function is used to optimize the parameters.

[0048] Further, the Pro-Clash method is specifically as follows: The conflicting entity sets are separated pairwise, then the logits of their corresponding segments are compared, and the candidate segment with the larger logits is retained.

[0049] In summary, the present invention has the following beneficial effects: The present application proposes a nested entity recognition model that fuses cross-position encoding and Dropup. First, the cross-position encoding is used to strengthen the distance representation between characters, improve the model's attention to word relationships, and make the entity boundaries clearer. Secondly, the Dropup method is used to confuse the features of the position information and character embeddings for more effective feature fusion, improve the robustness and generalization of the model. Finally, a decoding framework is used to identify the start and end of entities and solve the conflict problem, and entities are extracted more comprehensively. The results show that the CrossDp-NER model only models position features and performs partial interactions, without additional inputs such as vocabulary and pinyin, and its performance has been improved compared with the current mainstream models, indicating that position features are a key information in the entity recognition task, and at the same time, feature fusion is also an effective way to improve the effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the overall architecture diagram of the hierarchical pointer nested entity recognition model based on cross-attention in the embodiment of the present invention;

[0051] Figure 2 is the pointer network diagram of the feature interaction layer in the embodiment of the present invention;

[0052] Figure 3 is the nested pointer network diagram of the feature interaction layer in the embodiment of the present invention;

[0053] Figure 4 is the decoding flow chart of the decoding layer in the embodiment of the present invention;

[0054] Figure 5 is the entity conflict phenomenon diagram of the decoding layer in the embodiment of the present invention;

[0055] Figure 6 is the Pro-Clash comparison diagram in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following further Figures 1-6 describes the present invention in detail with reference to the attached

[0057] Embodiment: A hierarchical pointer nested entity recognition model based on cross-attention, such as Figure 1As shown, it includes embedding layer, feature interaction layer, decoding layer and prediction layer;

[0058] The embedding layer uses RoBERTa as the character encoding layer;

[0059] BERT is an encoding representation model built by Google using bidirectional Transformers. It has excellent performance in 11 NLP tasks and can be said to be a milestone. It can capture the contextual information of the sequence, comprehensively and diversely represent the text, and better understand the semantics and context. Therefore, it has become a basic step in the named entity recognition task.

[0060] The feature interaction layer injects position information and confuses features into the output of the embedding layer through cross position encoding and Drop-up feature fusion method, thereby strengthening feature representation.

[0061] Rotational Position Encoding (RoPE) is a position encoding model proposed by Su Jianlin. It encodes the position information to each dimension by rotation, so that the model can capture the relative position information between characters in the sequence and has good extrapolation properties. It has been widely used in large models such as LLaMA and ChatGLM. This application is based on RoPE and improves it.

[0062] The decoding layer builds an entity category layer based on the pointer network and uses a multi-label classification loss function to optimize parameters;

[0063] This application builds an entity category layer based on the pointer network and uses a multi-label classification loss function to optimize parameters. It is compatible with nested and non-nested entities and solves the intra-group nesting problem that the pointer network could not handle in the past.

[0064] Based on the pointer network method, the head and tail tags are used to obtain candidate entities. If the position information is not considered, the head and tail of any two entities will be identified, which affects the entity boundary judgment. In this regard, the present application designs a cross position encoding and drop-up feature fusion method to inject position information and feature confusion into the output of the embedding layer to strengthen the feature representation.

[0065] The pointer network marks each span with a start and an end, indicating the entity start and end positions, respectively. Figure 2 As shown in the figure, the annotation idea of ​​the pointer network is that "中" is the start character of the ORG type, and "学" is the end character of the ORG type. The entity extraction can be completed by extracting the first and last tags of the same type. However, this has a disadvantage that nested entities cannot be extracted, that is, the character "中" can only select the type with the highest probability, so the CONT entity "中国" cannot be marked and cannot be recognized.

[0066] Based on the above description, the entity category layer is designed based on the Start and End tags, such as Figure 3 As shown, from Figure 3 From the figure, we can see the increase of entity category layer, so that the CONT entity "China" and the ORG entity "Renmin University of China" can be predicted separately, realizing nested and non-nested entity extraction.

[0067] This part adopts a multi-label classification loss function to specifically compare the scores of the target category and the non-target category, so that the score of the target category is greater than that of the non-target category, solving the problem of category imbalance.

[0068] The prediction layer filters out the position labels less than 0 for the logits obtained by the decoding layer, and then traverses and collects the complete entity set, such as Figure 4 As shown in the figure, first find the character "中" greater than 0 in the Start tag, and then extract the candidate entities of "中国" and "中国的人人" according to the characters "国" and "民" greater than 0 in the End tag of its start tag. It can be seen that the extracted entities cover all, so there will be an entity conflict phenomenon, such as Figure 5 As shown, Figure 5 According to the above method, the "Shanghai" CONT entity and the "Haishihong" CONT entity will be extracted. It is obvious that the two entities intersect each other, and the red mark in the middle is the conflicting part. However, this situation is not allowed in actual tasks. In this regard, this application is inspired by CNN-NER and designs a processing method Pro-Clash, which separates the conflicting entity sets into pairs, and then compares the logits size of each corresponding segment, and retains the candidate segment with large logits to achieve the purpose of solving the problem.

[0069] by Figure 5 Take this as an example, the specific steps are as follows:

[0070] (1) Extract the two candidate entities "Shanghai" and "Haishihong";

[0071] (2) Add the logits of the first and last tags to get the logits of the entity; the logits of Shanghai is 2, and the logits of Haishihong is 0.5

[0072] (3) Compare the logits and select the entity with the largest logits, i.e. "Shanghai"

[0073] The Pro-Clash method can be used to further screen conflicting entities and retain the correct entities to the greatest extent possible.

[0074] Embodiment 2: A hierarchical pointer nested entity recognition method based on cross attention, characterized by comprising the following steps:

[0075] S1. Inject positional information into the character vector through cross-position encoding to obtain the relative position vector;

[0076] S2. Confuse the relative position vector and the initial character embedding through Drop-up feature confusion to strengthen the feature representation;

[0077] S3. Improve the pointer network structure, use the start and end classification matrices to mark the candidate entity set, and introduce the multi-label cross-entropy loss function to quickly train and reduce error propagation;

[0078] S4. Adopt the Pro-Clash method to filter the entities with interactions in the candidate intervals and further solve the entity conflict problem.

[0079] Furthermore, the character vector is obtained by inputting the characters and their corresponding positions in the text sequence into the RoBERTa model for encoding and representation;

[0080] The expression of the text sequence is as follows:

[0081] S = {x1, x2, x3... x n} (1)

[0082] where x n represents the characters in the text sequence, and n represents the corresponding position;

[0083] The formula of the RoBERTa is as follows:

[0084] RoBERTa_out = RoBERTa(S) (2)

[0085] As a preferred solution of this embodiment, the cross-position encoding specifically includes the following steps:

[0086] ① Implement the goal of the RoPE model, that is, include the distance information while taking the inner product of two vectors. The formula is as follows:

[0087] <f q (q m , m), f k (k n , n)> = g(q m , k n , m - n) (3)

[0088] where m - n represents the distance information, and q m , k n represent the query vector and the key vector at positions m and n;

[0089] ② Add the distance information of m - n to the inner product between positions m and n. The formula is as follows:

[0090] f q (q m ,m) = q m e imθ = (W q x m )e imθ

[0091] f k (k n ,n) = k n e inθ = (W k x n )e inθ (4)

[0092] g(q m ,k n ,m - n) = Re[q m k n * e i(m-n)θ = Re[(W q x m )(W k x n ) * e i(m-n)θ

[0093] In the formula, e imθ and e inθ are complex exponential functions represented by Euler's formula, which are related to trigonometric functions and expanded into rotation matrices, where i is the imaginary unit and e is the base of the natural logarithm; k n * is the conjugate complex number of k n , and Re represents taking the real part;

[0094] Among them, the expression for the value range of θ is as follows:

[0095] θ = [θ0,..., θ d / 2-1 (5)

[0096] In the formula, d is the dimensional length of the query vector, which is used for calculation by substituting into the sine and cosine functions;

[0097] ③ Perform rotational position encoding on the embedding vector and perform the reverse operation on the query vector at the same time. The formula is as follows:

[0098] q m = contrary(RoBERTa_out)·W Q

[0099] k n = RoBERTa_out·W K (6)​

[0100] V = RoBERTa_out · W V

[0101] Wherein, RoBERTa_out is the embedding vector, and contrary is the reverse operation to reverse-encode the query vector;

[0102] ④ Use the positional encoding in RoPE to multiply the query and key vectors by the rotation matrix for calculation. The formula is as follows:

[0103]

[0104] ⑤ Perform dot-product attention calculation on the query and key vectors after positional encoding and the value vector to obtain the final positional feature CrossPE_out. The formula is as follows:

[0105]

[0106] As a preferred solution of this embodiment, in step S2, the formula for Drop-up feature confusion is as follows:

[0107]

[0108] Wherein, respectively represent the embedding vector RoBERTa_out and the positional feature CrossPE_out. Dropout = 1 means that the encoding vector t1 of character i is set to 0 using Dropout, otherwise the t2 vector at the corresponding position is output.

[0109] Furthermore, in step S3, the method for improving the pointer network structure is specifically as follows: Build an entity category layer on the basis of the pointer network, and use a multi-label classification loss function to optimize the parameters. The formula is expressed as:

[0110]

[0111] Furthermore, the Pro-Clash method is specifically as follows: Separate the entity sets with conflicts in pairs, then compare the logits of their respective corresponding segments, and retain the candidate segment with the larger logits.

[0112] Example 3: Experimental verification:

[0113] To verify the effectiveness of the model, this embodiment selects nested and flat structure datasets. The nested dataset is CNERTA; the flat datasets are Resume, Weibo, and E-commerce.

[0114] The datasets are introduced in detail as follows:

[0115] (1) The CNERTA training set has 34,102 data, the validation set has 4,440 data, and the test set has 4,445 data; the proportions of nested entities are 31.25%, 29.50%, and 38.25% respectively;

[0116] (2) The Resume training set has 3,821 data, the validation set has 463 data, and the test set has 477 data;

[0117] (3) The Weibo training set has 1,350 data, 270 data, and 270 data;

[0118] (4) The E-commerce training set has 3,989 data, 500 data, and 500 data.

[0119] Baseline model: To verify the effectiveness of the model construction method in this embodiment, this embodiment selects relevant mainstream models for verification recently, as follows: Baseline, GlobalPointer, W2NER, SwM, RDKVC, BIFT, LEBERT, UnifiedNER, NEZHA-BC, HTLR, MLDF, Chinese nested named entity recognition model based on position embedding and multi-level prediction, BS, two-stage network and prompt learning, Chinese named entity recognition combining gazetteers and syntactic dependency trees.

[0120] Evaluation metrics: The commonly used evaluation criteria F1 value (F-measure), P value (accuracy), and R value (recall rate) are adopted. The calculation formulas are: P = RN / PN * 100%, R = RN / GN * 100%, F1 = 2PR / (P + R); where RN represents the number of correct entities in the prediction result; PN represents the total number of predicted entities; GN represents the number of correct entities in the test data.

[0121] Parameter settings: The model parameters are set as follows. CrossPE hidden layer: 768, learning rate: 5e-5, Drop-up: 0.3, Droupout: 0.1, RoBerta hidden layer: 768, optimizer: AdamW, Python version is 3.8, Pytorch version is 1.10.0, GPU is P100-16G, and the RoBerta model uses RoBerta_zh_12.

[0122] Considering the inconsistency in the number of data sets, the epoch is set to 20 on the Resume data set; the epoch is set to 30 on the Weibo data set, the epoch is set to 2 on the CNERTA data set, and the epoch is set to 20 on the E-commerce data set.

[0123] Comparative experiments: The experimental results on the Resume, Weibo, E-commerce, and CNERTA datasets are shown in Tables 1-4 below:

[0124] Table 1 Experimental Results on Resume

[0125]

[0126] Table 2 Experimental Results on Weibo

[0127]

[0128] Table 3 Experimental Results on E-commerce

[0129]

[0130] Table 4 Experimental Results on CNERTA

[0131]

[0132]

[0133] As can be seen from Tables 1-4 above:

[0134] (1) The F1 of the model in this paper on the Resume dataset is improved by 0.2%, 0.19%, 0.77%, 0.52%, 0.69%, 0.71%, 0.62%, 0.22%, and 0.22% respectively compared with W2NER, BS, LEBERT, HTLR, SwM, RDKVC, BIFT, MLDF, and the two-stage learning model;

[0135] (2) The F1 of the model in this paper on the Weibo dataset is improved by 1.16%, 1.37%, 1.66%, 0.09%, and 0.69% respectively compared with BS, LEBERT, W2NER, MLDF, and SwM, RDKVC;

[0136] (3) The F1 of the model in this paper on the E-commerce dataset is improved by 0.6%, 2.95%, 0.46%, and 7.83% respectively compared with Baseline, Combination, NEZHA-BC, and UnifiedNER;

[0137] (4) The F1 of the model in this paper on the CNERTA dataset is improved by 0.03%, 3.6%, 3.67%, and 16.04% respectively compared with Baseline, GlobalPointer, Effi-GlobalPointer, and position-based;

[0138] (5) On the Weibo and Resume common public datasets, compared with the three methods of introducing external information to construct span representation enhanced features, namely RDKVC, HTLR, and SwM, the structure of this paper is simpler and more effective. It only models the positions of characters and confuses features, and at the same time can achieve better performance.

[0139] Based on the above F1 value results, it shows that the performance of the method of this application is better than that of the current mainstream models, and it also proves that the structure of this application can learn the position distance representation between entities more fully. Moreover, the R value is also the best on the three datasets of Resume, Weibo, and CNERTA, indicating that the structure of this paper can effectively expand the entity boundaries and extract more and more correct entities.

[0140] Ablation experiment:

[0141] To verify the effectiveness of the structure proposed in this paper, ablation experiments were carried out on the above four datasets with the same parameters. The comparison models are as follows:

[0142] (1) Baseline: Use the decoding framework of this paper and remove the cross-position encoding and Dropup methods;

[0143] (2) Baseline + LLM-Dropup: Use the attention calculation for the RoPE encoding in the LLM and use the Dropup feature fusion method;

[0144] (3) CrossDp-NER: The method proposed in this paper.

[0145] Note: Since Dropup requires two types of feature inputs, it is used together with the position encoding

[0146] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications without creative contributions to this embodiment according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

[0147] Table 5 Results of Resume ablation experiment

[0148]

[0149] Table 6 Results of Weibo ablation experiment

[0150]

[0151] Table 7 Results of E-commerce ablation experiment

[0152]

[0153] Table 8 CNERTA Ablation Experiment Results

[0154]

[0155] As can be seen from the above Tables 5 - 8: In Resume, compared with the baseline, the F1 values of Baseline+LLM-Dropup and CrossDp-NER increase by 0.26% and 0.62% respectively; in Weibo, compared with the baseline, the F1 values of Baseline+LLM-Dropup and CrossDp-NER increase by 1.16% and 2.54% respectively; in E-commerce, compared with the baseline, the F1 values of Baseline+LLM-Dropup and CrossDp-NER increase by 0.65% and 0.6% respectively; in CNERTA, compared with the baseline, the F1 values of Baseline+LLM-Dropup and CrossDp-NER decrease by 0.1% and 0.05% respectively. The above changes in F1 values indicate that the Cross_PE and Dropup methods proposed in this paper can effectively improve the model performance. In addition, it can also be known that the method proposed in the Weibo dataset has the largest improvement, indicating that on datasets with low data volume and sparse entities, the structure of this application can maximize the capture of latent information between characters.

[0156] Pro-Clash Experiment: The above results prove that the improvement is most obvious on the Weibo dataset. Therefore, in this part, to verify the impact of the Pro-clash method on the model, a comparison is made on the Weibo dataset, as shown in Table 9 below and Figure 5 as shown.

[0157] Table 9 Resume Ablation Experiment Results

[0158] Table 9 Resume ablation experimental results

[0159]

[0160] The parameter settings of Pro-Class in Table 9 above are 1, 2, and 3, which means processing the first, the first two, and the first three in the entity conflict set. Because the number of conflicting entities in the prediction results takes a long time to process, for efficiency considerations, it is set to a smaller number.

[0161] It can be learned from the results that the F1 value of the Pro-Clash method reaches 75.61 when the parameter is 1. In addition, it can be seen that as the parameter increases, the P value continues to grow and the R value decreases slightly, indicating that the Pro-Clash method can effectively reduce the generation of incorrect entities, improve the accuracy rate, balance the recall rate, and achieve an improvement in the F1 value.

[0162] In summary, this application proposes a nested entity recognition model that combines cross-position encoding and Dropup. First, cross-position encoding is used to strengthen the distance representation between characters, improve the model's attention to word relationships, and make entity boundaries clearer. Second, the Dropup method is used to confuse the features of the position information and character embeddings to more effectively achieve feature fusion, improve the robustness and generalization of the model. Finally, a decoding framework is used to identify the beginning and end of entities and solve conflict problems, and entities are extracted more comprehensively. The results show that the CrossDp-NER model only models position features and performs partial interactions, without the need for additional inputs such as vocabulary and pinyin, and its performance has a certain improvement compared with current mainstream models, indicating that position features are a key information in the entity recognition task, and at the same time, feature fusion is also an effective way to improve the effect.

[0163] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications to this embodiment without creative contributions according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A hierarchical pointer nested entity recognition model based on cross attention, characterized by: Includes embedding layer, feature interaction layer, decoding layer and prediction layer; The embedding layer uses RoBERTa as the character encoding layer; The feature interaction layer injects position information and confuses features into the output of the embedding layer through cross position encoding and Drop-up feature fusion method, thereby strengthening feature representation. The decoding layer builds an entity category layer based on the pointer network and uses a multi-label classification loss function to optimize parameters; The prediction layer filters out position labels less than 0 for the logits obtained by the decoding layer, and then traverses and collects the complete entity set.

2. A hierarchical pointer nested entity recognition method based on cross attention, characterized in that: The following steps are involved: S1, inject position information into the character vector through cross position coding to obtain the relative position vector; S2, confuse the relative position vector and the initial character embedding through Drop-up feature confusion to enhance feature representation; S3. Improve the pointer network structure, use the head-tail classification matrix to mark the candidate entity set, and introduce the multi-label cross-loss entropy function to quickly train and reduce error propagation; S4. Pro-Clash method is used to filter entities that interact in candidate intervals to further solve the conflict problem between entities.

3. The hierarchical pointer nested entity recognition method based on cross attention according to claim 2 is characterized in that: In step S1, the character vector is obtained by inputting the characters in the text sequence and their corresponding positions into RoBERTa for encoding representation; The expression of the text sequence is as follows: S={x1,x2,x3...x n } (1) In the formula, x n Represents a character in a text sequence, and n represents the corresponding position; The formula of RoBERTa is as follows: RoBERTa_out=RoBERTa(S) (2).

4. The hierarchical pointer nested entity recognition method based on cross attention according to claim 2 is characterized in that: In step S1, cross position coding specifically includes the following steps: ① The RoPE model is used to achieve the goal of taking the inner product of two vectors while including distance information. The formula is as follows: <f q (q m ,m),f k (k n ,n)>=g(q m ,k n ,m-n) (3) In the formula, mn represents distance information, q m , k n Represents the query vector and key vector at positions m and n; ② Add the distance information of mn to the inner product between positions m and n. The formula is as follows: In the formula, e imθ 、e inθ Euler's formula represents the complex exponential function and trigonometric function, which is expanded into a rotation matrix, where i is the imaginary unit and e is the base of the natural logarithm; k n * k n The conjugate complex number of , Re represents the real part; The expression of the value range of θ is as follows: θ=[θ0,...,θ d / 2-1 ] (5) Where d is the dimension length of the query vector, which is substituted into the sine and cosine functions for calculation; ③ Rotate the embedding vector to encode the position, and perform the reverse operation on the query vector. The formula is as follows: Where RoBERTa_out is the embedding vector, contrary is the reverse operation to reversely encode the query vector; ④ Use the position encoding in RoPE to multiply the query and key vectors by the rotation matrix. The formula is as follows: ⑤ Perform dot product attention calculation on the position-encoded query and key vectors and the value vector to obtain the final position feature CrossPE_out. The formula is as follows:

5. The hierarchical pointer nested entity recognition method based on cross attention according to claim 2 is characterized in that: In step S2, the formula for Drop-up feature confusion is as follows: In the formula, They represent the embedded vector RoBERTa_out and the position feature CrossPE_out respectively. Dropout=1 means that the encoding vector t1 of character i is set to 0 using Dropout, otherwise the t2 vector of the corresponding position is used for output.

6. The method for hierarchical pointer nested entity recognition based on cross attention according to claim 2 is characterized in that: In step S3, the method for improving the pointer network structure is specifically: building an entity category layer based on the pointer network, and using a multi-label classification loss function to optimize parameters.

7. The method for hierarchical pointer nested entity recognition based on cross attention according to claim 2 is characterized in that: The Pro-Clash method is as follows: the conflicting entity sets are separated in pairs, and then the logits sizes of their corresponding fragments are compared, and the candidate fragments with large logits are retained.

Citation Information

Patent Citations

  • Named entity recognition method for contradictory mediation text based on efficient pointer network

    CN117010397A

  • Chinese predicate recognition method based on POS fusion feature and entity boundary diagnosis

    CN117252199A

  • Chinese named entity recognition method based on stacked grid structure information enhancement

    CN119227685A

  • System and method for natural language processing with pretrained language models

    US20220237378A1