Relationship enhanced named entity recognition method and device, medium and program product

Through the relationally enhanced named entity recognition method, BERT, convolutional neural network and Transformer model are used, combined with the conditional random field layer, the problem of difficulty in identifying discontinuous entities and composite entities is solved, and a higher recognition accuracy and completeness are achieved.

CN120106068APending Publication Date: 2025-06-06QIZHI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510197392.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional named entity recognition methods are difficult to accurately identify discontinuous entities and composite entities, resulting in insufficient accuracy and completeness of the recognition results.

Method used

The relationship-enhanced named entity recognition method is adopted to extract features through the BERT model, combine convolutional neural network and Transformer model to perform feature fusion and long-distance dependency modeling, generate a growth entity structure model, and optimize label sequences through the conditional random field layer.

Benefits of technology

It improves the accuracy and completeness of named entity recognition, can identify complex entities more accurately, and enhances the model's adaptability and robustness to different texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106068A_ABST
    Figure CN120106068A_ABST
Patent Text Reader

Abstract

The invention discloses a relation enhancement type named entity recognition method and device, a medium and a program product, and relates to the field of data processing. The method comprises the following steps: performing feature extraction on text data through a BERT model to obtain a plurality of tokens and semantic feature vector sequences; performing different-scale convolution operations by using a convolutional neural network to obtain local character combination features, and fusing the local character combination features with the semantic feature vector sequence to form a feature enhanced semantic feature vector sequence; the Transform model determines a long-distance dependency relationship of semantic feature vectors in the sequence by using an attention mechanism; generating a long entity structure model based on the long-distance dependency relationship, and then determining an enhanced semantic feature vector sequence; and performing tag sequence optimization processing on the named entity through a conditional random field layer, determining sequence labeling loss, boundary detection loss and relation prediction loss, determining total loss, and adjusting a tag sequence so as to determine the named entity. According to the method, the accuracy and integrity of named entity recognition can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a relationship-enhanced named entity recognition method, device, medium and program product. Background Art

[0002] Named Entity Recognition (NER) is an important category in the field of NLP (natural language processing). It uses the model's understanding of natural language to classify and extract characters in the target text (people, time, place, products, services, works, etc.).

[0003] Currently, the NER task mainly uses the BIO annotation rule, namely "begin, end, other" (B-product, I-product, O). Although this BIO annotation is simple and efficient and can meet most application scenarios, it can only annotate simple continuous entities (such as "refrigerator" and "beef jerky"), and cannot accurately, rigorously and completely annotate non-continuous entities and composite entities (such as "large generator" and "medium generator" in "large, medium and small generators").

[0004] This results in the inability of traditional small text generation models (such as BERT) to effectively learn such entities, and thus cannot accurately identify and output them in predictions, which limits the performance of NER tasks and prevents the accuracy and completeness of the model's output results from being further improved. Summary of the invention

[0005] In view of the above-mentioned technical problems and defects, the object of the present invention is to provide a relationship-enhanced named entity recognition method, device, medium and program product, which can effectively improve the accuracy and completeness of named entity recognition.

[0006] To achieve the above objectives, in a first aspect, the present invention provides a relationship-enhanced named entity recognition method, comprising: Perform feature extraction on the text data to be processed through the BERT model to obtain multiple tokens and a semantic feature vector sequence composed of the semantic feature vectors of the multiple tokens; The convolution neural network is used to perform convolution operations of different scales on the semantic feature vector sequence to obtain local character combination features; The local character combination feature is fused with the semantic feature vector sequence to obtain a feature-enhanced semantic feature vector sequence; The Transformer model uses the attention mechanism to determine the long-distance dependency of the semantic feature vector in the feature-enhanced semantic feature vector sequence. The long-distance dependency is used to characterize the positional relationship and semantic relationship between the tokens. generating a long entity structure model based on the long-distance dependency; Determine a relation-enhanced semantic feature vector sequence according to the long entity structure model; Through the conditional random field layer, the label sequence of the relation-enhanced semantic feature vector sequence is optimized to obtain an optimized label sequence; Determine the sequence labeling loss, boundary detection loss, and relationship prediction loss for optimizing the label sequence; the sequence labeling loss is used to measure the difference between the predicted label sequence and the true label sequence output by the overall model; the boundary detection loss is used to evaluate the deviation of the accuracy of named entity boundary positioning from the conditional random field layer to the overall model; the relationship prediction loss is used to quantify the difference between the semantic relationship prediction and the actual relationship between the tokens contained in the potential named entity from the conditional random field layer to the overall model; The total loss is determined based on the sequence labeling loss, boundary detection loss, relationship prediction loss and the total loss function; the expression of the total loss function includes: ; Among them, L total represents the total loss, L seq represents the sequence labeling loss, L boundary represents the boundary detection loss, L relation represents the relationship prediction loss, α, β, γ are weight coefficients, λ 1 , 2 ,,λ 3 are normalization coefficients, δ is the nonlinear gating parameter, η, θ, and μ are the nonlinear intensity control parameters; Adjust the label sequence optimization processing and the overall behavior of the model according to the total loss; The named entity corresponding to the text data is determined according to the adjusted optimized label sequence.

[0007] The beneficial effects of the present invention include: the above method can effectively improve the accuracy and completeness of named entity recognition. The BERT model is used to extract features, which can accurately capture the semantics of the text and adapt to various text types. The convolutional neural network performs convolution operations of different scales to obtain local character combination features, enhances the perception of local semantic changes, and provides richer information for subsequent processing after fusion with the semantic feature vector sequence. The attention mechanism of the Transformer model determines long-distance dependencies, effectively captures the position and semantic relationship between tokens, and is particularly conducive to processing named entity recognition in long texts. A long entity structure model is generated based on long-distance dependencies, which can more accurately identify long entity names. The conditional random field layer optimizes the label sequence to ensure that the named entity recognition results are accurate and reliable, reduce erroneous annotations, and improve the overall recognition accuracy and completeness. At the same time, the sequence annotation, boundary detection and relationship prediction losses are determined, and then the total loss is obtained. The model is adjusted according to the total loss, and all aspects of deviations are comprehensively considered to improve the accuracy, completeness and robustness of named entity recognition, and enhance the adaptability of the model to different texts.

[0008] In a second aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation method of the first aspect or the second aspect.

[0009] In a third aspect, the present invention provides a computer-readable storage medium comprising instructions, which, when executed on the electronic device, causes the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation of the first aspect or the second aspect.

[0010] In a fourth aspect, the present invention provides a computer program product comprising instructions, which, when the computer program product is run on the electronic device, enables the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation of the first aspect or the second aspect.

[0011] It is understandable that the electronic device provided in the second aspect, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided in the present invention. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a schematic diagram of the structure of a NER model architecture according to an embodiment of the present invention; Figure 2 is a flow chart of a relationship-enhanced named entity recognition method according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the architecture of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The terms used in the following embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to be limiting of the present invention. As used in the specification of the present invention, the singular expressions "a", "a", "above", "the" and "this" are intended to also include plural expressions, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present invention refers to any or all possible combinations comprising one or more of the listed items.

[0014] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood as implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.

[0015] The embodiment of the present invention provides a NER model architecture with multi-level feature extraction, such as Figure 1 As shown in FIG, the NER model architecture includes a feature extraction layer, a feature enhancement layer, a relationship modeling layer, and a CRF (Conditional Random Field) layer.

[0016] Among them, the feature extraction layer uses the powerful semantic understanding ability of the BERT pre-trained model to map each token in the input processed text data into a vector representation containing rich contextual semantic information, providing basic semantic features for the entire model.

[0017] The BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model based on a deep bidirectional Transformer structure. It can extract rich semantic feature vectors for each token (the smallest semantic unit, including words or characters, etc.) in the text to be processed. In this process, BERT uses its knowledge and parameters pre-trained on a large-scale corpus to deeply understand the input text. For example, for the sentence "The fruit company has released a new mobile phone", BERT will generate a specific semantic feature vector for each token such as "fruit company", "release", "released", "new model", "mobile phone", etc. This vector contains the semantic information of the token in the current context, such as part of speech, semantic role, etc. In this way, the feature extraction layer provides a basic semantic representation for subsequent processing steps, laying an important foundation for the entire named entity recognition task.

[0018] The feature enhancement layer uses a convolutional neural network (CNN) to perform convolution operations of different scales on the semantic feature vector sequence from the feature extraction layer. CNN can effectively extract local character combination features, which is critical for named entity recognition tasks. For example, when processing the sentence "The fruit company has released a new mobile phone", CNN can extract features of local character combinations such as "fruit", "company", "new model", and "mobile phone", which can capture local semantic patterns and grammatical structures in the text. Then, these local character combination features are fused with the semantic feature vector sequence from the feature extraction layer to obtain a feature-enhanced semantic feature vector sequence. This fusion operation further enriches the representation of semantic features, allowing the model to better understand the semantic information in the text and enhance the recognition ability of named entities.

[0019] The relational modeling layer uses the Transformer model to determine the long-distance dependency of the semantic feature vectors in the feature-enhanced semantic feature vector sequence through the attention mechanism. This long-distance dependency can characterize the positional relationship and semantic relationship between tokens, which is crucial for identifying complex entities. For example, when processing the sentence "The fruit company has released a new mobile phone, which has attracted widespread attention from consumers", the relational modeling layer can capture the semantic association between "fruit company" and "mobile phone" and their positional relationship in the sentence through the attention mechanism. Based on this long-distance dependency, the relational modeling layer can build a long entity structure model to further improve the recognition ability of complex named entities. In addition, the relational modeling layer can also enhance the modeling ability of the positional relationship and semantic relationship between characters through the position-aware layer and the adaptive weight balancing mechanism, so that the model can more accurately capture the relationship between tokens and improve the accuracy of named entity recognition.

[0020] The Transformer model of this embodiment is an improved Transformer model and has the following characteristics: 1. Local perception mechanism: In natural language processing, although the traditional Transformer structure performs well in many tasks, it is not the best choice for named entity recognition tasks in specific scenarios such as business keyword extraction. The local perception mechanism in this embodiment adopts a local multi-head attention mechanism based on a sliding window, which makes an innovative improvement on this problem.

[0021] This local multi-head attention mechanism sets a sliding window on the input information, allowing the model to focus more on the semantic associations of the adjacent areas of entity words. For example, when processing text related to the main business words of a company, for the sentence "Smart home appliance manufacturers focus on the production of high-end refrigerators and washing machines", the local multi-head attention mechanism will focus more on entity words such as "smart home appliances", "refrigerators", "washing machines" and their adjacent areas, such as the association between "manufacturing companies" and "smart home appliances", and the relationship between "high-end" and "refrigerators". Such a design significantly improves the model's ability to understand local contexts, because in scenarios such as extracting the main business words of a company, key information is often concentrated in specific entity words and their surrounding semantic environments. Compared with the traditional global self-attention mechanism, the local perception mechanism avoids indiscriminate attention to the entire input information, thereby more efficiently extracting local semantic features related to specific tasks, providing a more accurate information basis for accurately identifying named entities such as the main business words of a company.

[0022] 2. Entity-aware attention: In the named entity recognition task, traditional methods often face challenges, especially for the recognition of non-continuous entities. The entity-aware attention mechanism effectively solves this problem by introducing entity position awareness and dynamically expanding the attention window according to the recognized entity position.

[0023] For example, in the text "A new smartphone software developed by a certain technology company has received widespread attention in the market", when the model recognizes the entity "smartphone software", the entity-aware attention mechanism will dynamically expand the attention window according to its position. In this way, the model can not only pay attention to the entity "smartphone software" itself, but also capture a wider range of contextual information around it, such as the R&D relationship between "a certain technology company" and the entity, and the market performance of the entity reflected by "receiving widespread attention in the market". This mechanism greatly improves the ability to recognize non-continuous entities, because non-continuous entities may be scattered in the text. By dynamically expanding the attention window, these scattered parts can be better associated, thereby accurately identifying complex non-continuous entities.

[0024] 3. Location Awareness Layer: Through independent position query (pos_query) and position key (pos_key) matrices, the model's ability to model the position relationship between characters (tokens) in the text is enhanced. In the named entity recognition task, especially for the internal structure recognition of non-continuous entities, the position relationship is crucial.

[0025] For example, when processing the sentence "A large industrial automation equipment manufacturer has launched a new intelligent control system", the pos_query and pos_key matrices of the position-aware layer can help the model clarify the positional relationship between each token. For example, the positional relationship between "large" and "industrial automation equipment", the order of "manufacturer" and "launch", etc. In this way, the position-aware layer can provide the model with clear information about the position between characters, allowing the model to better understand the structure and semantics of the text. For non-continuous entities, such as "large-scale industrial automation equipment" and "new intelligent control system", the position-aware layer can help the model determine the positional relationship between the internal parts of these entities, so as to more accurately identify and understand the internal structure of non-continuous entities, and provide more refined position clues for named entity recognition tasks.

[0026] 4. Adaptive weight balancing: In different input data text environments, the importance of token position information and semantic relationships often varies. In order to improve the flexibility and adaptability of the model, the present invention introduces learnable position weight and relationship weight parameters to achieve adaptive weight balance.

[0027] For example, when processing technical document texts, position information may be relatively more important, because technical documents usually have stricter structural and sequential requirements. At this time, the model can automatically increase the parameter value of the position weight through learning, so that the model pays more attention to the positional relationship between characters. When processing literary works or daily conversation texts, semantic relationships may be more critical, and the model will adjust the relationship weight parameters accordingly, paying more attention to the semantic connection between tokens. This adaptive weight balancing mechanism enables the model to automatically adjust the relative importance of position information and semantic relationships according to the different characteristics of the input data text, so that named entity recognition can be better performed in various text environments, improving the generalization and adaptability of the model.

[0028] The CRF layer performs label sequence optimization on the relationship-enhanced semantic feature vector sequence from the relationship modeling layer. The CRF layer can consider the transition probability between labels, thereby optimizing the label sequence output by the model to ensure that the output label sequence is more reasonable in terms of syntax and semantics. For example, in the named entity recognition task, the CRF layer can learn the transition probability between different named entity labels based on the training data, such as the probability of transferring from "B-PERSON" (indicating the beginning of the name entity) to "I-PERSON" (indicating the inside of the name entity) is high, while the probability of directly transferring from "B-PERSON" to "O" (indicating non-entity) is low. In the prediction stage, the CRF layer optimizes the label sequence of the relationship-enhanced semantic feature vector sequence based on these transition probabilities, so that the output optimized label sequence more accurately reflects the named entities in the text. In this way, the CRF layer improves the accuracy and completeness of the model's named entity recognition.

[0029] The embodiment of the present invention provides a relationship-enhanced named entity recognition method, which applies the NER model provided by the above embodiment. First, the text data is feature extracted by the BERT model to obtain multiple tokens and semantic feature vector sequences, providing basic semantic information for subsequent processing. Then, the semantic feature vector sequence is subjected to convolution operations of different scales by using a convolutional neural network to obtain local character combination features, thereby enhancing the understanding of local semantics. The local character combination features are fused with the semantic feature vector sequence to further enrich the feature representation. Then, the long-distance dependency relationship is determined by using the attention mechanism of the Transformer model. This relationship can accurately characterize the positional relationship and semantic relationship between tokens, thereby effectively processing non-continuous entities and composite entities, making up for the shortcomings of traditional labeling rules. The long entity structure model generated based on the long-distance dependency relationship provides strong support for the accurate identification of complex entities. Finally, the relationship-enhanced semantic feature vector sequence is optimized by the conditional random field layer to ensure that the output optimized label sequence can accurately correspond to the named entities in the text data, thereby improving the accuracy and completeness of the model output results.

[0030] Combine the following Figure 2 The method of this embodiment is specifically described, including the following steps: Step 201 : Perform feature extraction processing on the text data to be processed through the BERT model to obtain multiple tokens and a semantic feature vector sequence composed of the semantic feature vectors of the multiple tokens.

[0031] Specifically, in the feature extraction layer of the NER model, the BERT model is used to perform feature extraction on the text data to be processed. First, the input text data will be segmented according to its established segmentation rules. For example, for a text like "Today is sunny, suitable for going out to play", it will be segmented into tokens such as "today", "sunny", "suitable", "go out", "play" and so on according to the corresponding language rules and BERT's preset segmentation method.

[0032] Then, these tokens will be input into the BERT model in turn. The BERT model generates a corresponding semantic feature vector for each token based on its rich semantic knowledge learned from pre-training on large-scale corpus and the powerful representation ability of the deep bidirectional Transformer architecture. For example, the token "sunny" will be calculated through complex calculations within the model, taking into account its position in the entire text, context, and general semantics in the language, and output a specific feature vector containing rich semantic information such as part of speech, semantic role, and semantic association.

[0033] In this way, all tokens after word segmentation will be processed and generate their own semantic feature vectors. These semantic feature vectors are combined in sequence to form a semantic feature vector sequence composed of the semantic feature vectors of multiple tokens, providing basic and key semantic representation information for subsequent further text processing.

[0034] Step 202: Perform convolution operations of different scales on the semantic feature vector sequence through a convolutional neural network to obtain local character combination features.

[0035] Among them, the local character combination (n-gram) feature is a widely used text feature representation method in natural language processing. It captures local semantic and grammatical information based on n consecutive characters or words in the text. Specifically, for a given text sequence, the n-gram feature divides the text into a series of continuous subsequences according to a fixed length n. For example, when n=3, for the text "I love natural language processing", 3-gram subsequences such as "I love myself", "love nature", "natural language", "natural language", "language place", and "processing" can be obtained. These subsequences can reflect the local character combination patterns in the text. Different n values ​​can capture information of different granularities. When n is small, it focuses on capturing the local collocation relationship between characters or words, such as bigrams can find common collocations between adjacent words; when n is large, it can cover richer contextual information, such as trigrams and quadruplets can represent more complex local semantic structures.

[0036] In practical applications, n-gram features can be used for tasks such as text classification, information retrieval, and language models. By counting the frequency of n-grams in the text, they are converted into vector representations and then used as input features of the model to help the model learn the local semantics and grammatical rules in the text and improve its ability to understand and process the text.

[0037] In this step, the semantic feature vector sequence is input into the feature enhancement layer of the NER model for processing. When the semantic feature vector sequence is subjected to convolution operations of different scales through the convolutional neural network (CNN) to obtain local character combination features, the semantic feature vector sequence previously generated by the BERT model is first used as input data. The convolution kernel in the convolutional neural network will slide on this sequence, and convolution kernels of different scales can capture local information of different ranges.

[0038] For example, a smaller-scale convolution kernel can focus on the close relationship between several adjacent tokens, such as capturing the semantic relationship between "large", "medium" and "small" in "large, medium and small"; while a larger-scale convolution kernel can cover a wider area and obtain the features of longer local character combinations, such as understanding a slightly longer local semantic structure such as "large, medium and small generators".

[0039] During the convolution operation, the convolution kernel performs a dot product operation with the input semantic feature vector sequence, and then performs a nonlinear transformation through the activation function to extract various local feature patterns. As the convolution kernel continues to slide on the sequence, the local information of each position is extracted and integrated, and finally a rich local character combination feature is obtained. These features can better reflect the semantic characteristics of the local area in the text and provide more detailed information for subsequent feature fusion and entity recognition.

[0040] Step 203: fuse the local character combination feature with the semantic feature vector sequence to obtain a feature-enhanced semantic feature vector sequence.

[0041] Specifically, the feature enhancement layer aligns and adjusts the dimensions of the local character combination features and the semantic feature vector sequence to ensure that they are compatible and fused in shape and representation. For example, if the dimensions of the local character combination features are not completely consistent with the semantic feature vector sequence, some linear transformations or pooling operations may be needed to adjust them.

[0042] Then, a suitable fusion strategy is used to fuse the two. A common method is direct concatenation, which is to concatenate the local character combination features with the corresponding semantic feature vectors by dimension to form a new feature vector with higher dimensions, thereby combining the rich information of the local character combination with the overall semantic information of the original semantic feature vector. Another method is weighted summation, which assigns certain weights to the local character combination features and the semantic feature vector sequence respectively, and performs weighted summation according to their importance in the entity recognition task to obtain the fused feature vector.

[0043] Finally, the fused feature vectors are rearranged in the original order to form a feature-enhanced semantic feature vector sequence. This new sequence not only retains the global semantic information in the original semantic feature vector sequence, but also incorporates the local semantic details contained in the local character combination features, thus providing the model with a more comprehensive and richer semantic representation in subsequent processing, which helps to more accurately identify named entities.

[0044] Step 204, using the attention mechanism through the Transformer model, determine the long-distance dependency relationship of the semantic feature vector in the feature-enhanced semantic feature vector sequence.

[0045] The long-distance dependency is used to characterize the positional relationship and semantic relationship between the tokens.

[0046] The feature-enhanced semantic feature vector sequence is input to the relational modeling layer of the NER model, which uses the feature-enhanced semantic feature vector sequence as input through the Transformer model. The self-attention mechanism in the Transformer model calculates the degree of association between each semantic feature vector and all other semantic feature vectors in the sequence. For a specific semantic feature vector in the sequence, such as the vector corresponding to the token "wind power generation" in the text "large wind power generation equipment", the self-attention mechanism considers the mutual influence between this vector and other vectors, such as the vectors corresponding to tokens such as "large" and "equipment".

[0047] During the calculation process, position encoding will be introduced to help the model distinguish tokens at different positions, so as to better capture the positional relationship between tokens. For example, the model can use position encoding to determine that "large" comes before "wind power generation" and "equipment" comes after "wind power generation", thus establishing a preliminary relationship cognition based on location.

[0048] At the same time, the self-attention mechanism measures the semantic similarity and relevance through operations such as dot product operations and normalization between different vectors, thereby representing the semantic relationship between tokens. For example, since "wind power generation" has a semantic modification relationship with "large-scale" and a semantic whole-part relationship with "equipment", the self-attention mechanism will adjust the attention weight according to these semantic connections.

[0049] By continuously iteratively calculating and updating the attention weights, the Transformer model finally determines the long-distance dependencies between the semantic feature vectors in the feature-enhanced semantic feature vector sequence. This long-distance dependency not only contains the positional relationship information between tokens, which can clarify the order and relative position of different tokens in the text, but also covers the semantic relationship information, reflecting the semantic connection and logical relationship between tokens, providing rich contextual information and structural clues for subsequent named entity recognition tasks.

[0050] In some embodiments, step 204 may further include: determining the positional relationship between the tokens through a position awareness layer of the Transformer model.

[0051] The location perception layer includes a location query matrix and a location key matrix. The location query matrix is ​​used to determine the location of the token, and the location key matrix is ​​used to determine the relative position relationship between the tokens.

[0052] Specifically, first, when the feature-enhanced semantic feature vector sequence is input into the position-aware layer of the Transformer model, the position query matrix starts to operate. It will generate corresponding position representation information for each token in the sequence, as if each token is labeled with an exclusive "position label" to clearly determine the specific position of the token in the entire text sequence.

[0053] For example, when processing a token sequence converted from a text such as "Today is sunny and suitable for going out for fun", through the position query matrix, the token "today" has its corresponding clear position identifier, and it can be known that it is at the beginning of the sequence.

[0054] At the same time, the position key matrix is ​​also involved, which focuses on determining the relative position relationship between tokens. It interacts with the position features of different tokens, such as calculating the degree of correlation between the position vectors corresponding to two tokens, to clearly describe their relative positions such as order and distance.

[0055] Taking the previous text as an example, the position key matrix can make it clear that "sunshine" is in the subsequent position relative to "today" and can measure the distance between the two. By sorting out the relative position relationship between each token in this way, it can build an accurate understanding of the text structure for the model, allowing the model to better utilize this position relationship information in subsequent processing, integrate it with semantic information, and then more accurately grasp the meaning of the entire text and the relationship between the various parts, providing an important basic guarantee for tasks such as named entity recognition.

[0056] Step 205: Generate a long entity structure model based on the long-distance dependency relationship.

[0057] Specifically, the relational modeling layer first uses the determined long-distance dependencies to further analyze the complex relationships between the tokens corresponding to each semantic feature vector in the feature-enhanced semantic feature vector sequence. For long entity expressions containing multiple parts, such as "ultra-high voltage smart grid system", the long-distance dependencies can clarify the modification effect of "ultra-high voltage" on "smart grid system" and the specific semantic association between "smart" and "grid system".

[0058] Then, combining the position information and semantic relationships, a framework is constructed that can accurately describe the internal structure and external connections of long entities. For example, the order of "ultra-high voltage", "intelligent", and "grid system" is determined based on the positional relationship between tokens, and the modification, limitation and other relationships between them are determined based on semantic relationships. Then, this framework is continuously adjusted and optimized so that the model can better adapt to different types of long entity structures. By learning and training a large number of different long entity samples, the model gradually grasps the common structural patterns and change laws of various long entities.

[0059] Finally, a complete long entity structure model is formed. This long entity structure model can clearly present the hierarchical relationship between the components of the long entity, the semantic connection, and the interaction with the external text. In practical applications, when new text data is input, the long entity can be accurately identified and parsed based on this long entity structure model, providing strong support for named entity recognition tasks.

[0060] In this embodiment, the long entity structure model can regard the long entities in the text as a complex structure composed of multiple tokens with semantic associations. In the process of building the long entity structure model, not only the semantics of a single token is focused on, but also the positional relationship and semantic relationship between tokens are considered. By deeply analyzing a large amount of text data features, the model can adaptively determine the weight parameters of the positional relationship and semantic relationship therein. For example, for some long entities with clear semantic boundaries but relatively flexible positions, the weight of the semantic relationship may be increased; and for some long entities that need to be assisted by contextual position information, the weight of the positional relationship will be increased accordingly.

[0061] The long entity structure model has strong dynamic adaptability and flexibility. In practical applications, facing different fields, different types of texts and various complex and changeable long entity situations, it can automatically adjust the structure and parameters of the model according to the specific situation. For example, when processing long entities of professional terms in scientific and technological literature, the model will focus more on the in-depth mining and analysis of semantic relationships to accurately identify those long entities with specific domain meanings; when processing long entities such as institution names in news reports, it will comprehensively consider the location and semantic relationships to ensure the accurate definition of the scope and category of the entity.

[0062] The advantage of the long entity structure model is that it can effectively deal with long entities and complex entities that are difficult for traditional named entity recognition methods to handle. Taking the expression "large, medium and small generators" as an example, traditional methods may have inaccurate recognition problems due to the discontinuity and structural complexity of the entity. The long entity structure model can clearly identify the various sub-entities such as "large generators" and "medium-sized generators" and their relationships through the precise grasp of long-distance dependencies, thereby accurately and completely identifying the entire long entity. This greatly improves the accuracy and completeness of named entity recognition and provides a more reliable foundation for subsequent natural language processing tasks.

[0063] Step 206: Determine a relationship-enhanced semantic feature vector sequence according to the long entity structure model.

[0064] Specifically, first, we use the constructed long entity structure model as a reference framework to deeply analyze the key information such as the hierarchical relationship, semantic association, and position order between the components contained therein. For example, for a long entity such as "large-scale intelligent industrial automation control system" presented in the long entity structure model, it is clear that "large" is a modification of the whole, "intelligent" and "industrial automation control" have a synergistic semantic connection, and "system" points out the category to which it belongs, and knows the specific order between them.

[0065] Then, based on this relationship information, we return to the corresponding feature-enhanced semantic feature vector sequence and re-examine and adjust the weight of the tokens represented by each semantic feature vector. For the semantic feature vectors corresponding to the tokens that are in the core semantic position and play a key role in the long entity structure, their importance weights are appropriately increased; and for the vectors corresponding to the relatively minor modifying tokens, their weights are reasonably set.

[0066] Then, through specific integration and mapping operations, these semantic feature vectors that have been sorted out and adjusted in weight according to the long entity structure model are rearranged and combined in an order and manner that conforms to the internal logic and external associations of the long entity, thereby generating a new relationship-enhanced semantic feature vector sequence that is more in line with the characteristics of the long entity structure.

[0067] Among them, this relationship-enhanced semantic feature vector sequence can better reflect the overall semantic and structural characteristics of long entities, and provide more accurate and effective feature representation for subsequent further processing using conditional random field layers and accurately completing named entity recognition tasks.

[0068] Step 207 , performing label sequence optimization processing on the relationship-enhanced semantic feature vector sequence through a conditional random field layer to obtain an optimized label sequence.

[0069] First, the relationship-enhanced semantic feature vector sequence is used as the input of the CRF layer (conditional random field layer). The CRF layer will calculate an initial label score distribution for each feature vector in the input sequence based on its internal transfer matrix and emission matrix. The transfer matrix is ​​used to describe the transition probability between different labels. For example, in named entity recognition, the transition probability from "B-ORG" (indicating the beginning of the organizational entity) to "I-ORG" (indicating the middle part of the organizational entity) reflects the order and dependency between labels. The emission matrix is ​​used to calculate the emission probability of each feature vector corresponding to each label, that is, given a feature vector, the possibility that it belongs to different labels.

[0070] Then, a dynamic programming algorithm, such as the Viterbi algorithm, is used to find the most likely path among all possible label sequences, that is, the globally optimal label sequence. In this process, the CRF layer will comprehensively consider the label score of each position and the transition probability between labels to avoid unreasonable label combinations, such as "B-LOC I-PER" (a location entity suddenly converted to a person name entity) that does not conform to semantic logic.

[0071] Finally, the most likely label sequence found is output as the optimized label sequence. This optimized label sequence fully utilizes the modeling ability of the CRF layer for label transfer relations, more reasonably annotates the relation-enhanced semantic feature vector sequence, improves the accuracy and consistency of the labels, and makes the final named entity recognition results more reliable and accurate.

[0072] Step 208, determining the sequence labeling loss, boundary detection loss, and relationship prediction loss of the optimized label sequence.

[0073] Among them, the sequence labeling loss is used to measure the degree of difference between the predicted label sequence and the true label sequence of the overall output of the model. In sequence labeling tasks such as named entity recognition, the NER model needs to predict a corresponding label for each token in the input text, such as a person's name, a place name, an organization name, etc. The closer the predicted label sequence is to the true label sequence, the more accurate the model's prediction is. For example, for the sentence "Xiao Ming is studying at Peking University", the true label sequence may be "name-place-verb". If the label sequence predicted by the model is also "name-place-verb", then the sequence labeling loss is 0, indicating that the prediction is completely correct; if the prediction is "name-verb-place", there is a difference from the true label sequence, and the sequence labeling loss will be greater than 0. By calculating this degree of difference, you can intuitively understand the performance of the model in sequence labeling, and then adjust and optimize the model.

[0074] Boundary detection loss is used to evaluate the degree of deviation from the accuracy of named entity boundary positioning from the conditional random field layer to the model as a whole. In named entity recognition, it is crucial to accurately define the boundaries of named entities, such as determining that a person's name "Zhang San Li Si" is a complete entity, rather than mistakenly splitting it into two entities, "Zhang San" and "Li Si". The conditional random field layer plays a role in constraining and optimizing the boundaries of named entities in the model. If the positioning of the named entity boundaries by the model as a whole deviates from the actual boundaries, boundary detection loss will occur. For example, for a long compound word "smart home appliance products", if its boundaries are mistakenly determined as two parts, "smart home appliance" and "product", instead of being regarded as a whole named entity, boundary detection loss will occur. Boundary detection loss can reflect the accuracy of the model in boundary positioning by measuring the difference between the named entity boundaries predicted by the model and the actual boundaries. By optimizing the boundary detection loss, the model can be prompted to more accurately identify the boundaries of named entities and improve the performance of named entity recognition.

[0075] The relationship prediction loss is used to quantify the degree of difference between the semantic relationship prediction and the actual relationship between the tokens contained in the potential named entity from the conditional random field layer to the model as a whole. When processing named entities, it is necessary not only to identify the entity itself, but also to clarify the semantic relationship between the tokens within the entity, such as the modification relationship between "beautiful" and "flowers" in "beautiful flowers". The conditional random field layer will predict these semantic relationships in the model. If the predicted semantic relationship does not match the actual relationship, a relationship prediction loss will be generated. For example, for the entity "tall trees", if the model predicts that "tall" and "trees" are in a parallel relationship rather than a modification relationship, then the relationship prediction loss will reflect the degree of this prediction error. This loss can be used to adjust the model's ability to predict semantic relationships and improve the model's understanding and grasp of the semantics of named entities.

[0076] Step 209, determining the total loss according to the sequence labeling loss, the boundary detection loss, the relationship prediction loss and the total loss function.

[0077] Specifically, the sequence labeling loss, boundary detection loss, and relationship prediction loss are brought into the total loss function to calculate the total loss. The expression of the total loss function includes: ; Among them, L total represents the total loss, L seq represents the sequence labeling loss, L boundary represents the boundary detection loss, L relation represents the relationship prediction loss, α, β, γ are weight coefficients, λ 1 , 2 ,,λ3 are normalization coefficients, δ is a nonlinear gating parameter; η, θ, μ are nonlinear intensity control parameters, specifically power parameters.

[0078] Specifically, α is used to control the sequence labeling loss L seq The linear weight of determines the importance of the label prediction task.

[0079] β is used to control the boundary detection loss L boundary The linear weight of , which adjusts the priority of entity boundary positioning.

[0080] γ is used to control the relationship prediction loss L relation The linear weight of , affects the optimization strength of semantic relationship modeling.

[0081] δ is used to adjust the scaling of the tanh (hyperbolic tangent function) term in the denominator to balance the contribution ratio of linear and nonlinear.

[0082] η is the nonlinear power parameter of the sequence labeling loss, which is used to amplify (η>1η>1) or suppress (η<1) the impact of label errors through exponential operations; θ is the nonlinear power parameter of the boundary detection loss, which is used to adjust the nonlinear expression strength of the boundary positioning error; μ is the nonlinear power parameter of the relationship prediction loss, which is used to control the sensitivity of semantic relationship errors and support dynamic modeling of complex relationships; λ 1 , 2 ,,λ 3 They are used to maintain the dimensional consistency of each loss term and do not participate in training.

[0083] The above total loss function effectively improves the performance of named entity recognition (NER) models in complex scenarios through a dynamic nonlinear fusion mechanism. In practical applications, such as processing long terms in medical texts (such as "atypical antipsychotic dystonia") or non-standard expressions in social media (such as "the new xPhone released by the fruit company is amazing"), the model needs to take into account label accuracy (such as correctly labeling "fruit" as a brand), boundary accuracy (avoiding splitting "xPhone" into "x" and "Phone"), and semantic relationship rationality (understanding the relationship between "fruit" and "xPhone"). By introducing power parameters (η, θ, μ), the model can adaptively adjust the nonlinear response of each loss. For example, reduce the sensitivity of boundary loss in noisy text (θ<1), or amplify the weight of semantic relationships in nested entity scenarios (μ>1); while the gating parameter δ and the normalized tanh function ensure training stability and prevent gradient explosion.

[0084] This design enables the model to more flexibly balance multi-task conflicts and achieve higher-precision entity recognition in data with low resources, high noise, or complex entity structures (such as long clauses in legal contracts and nested terms in biomedical literature), while adding only a small number of parameters and taking into account the efficiency requirements of industrial-level deployment.

[0085] Step 210, adjust the label sequence optimization processing and the overall behavior performance of the model according to the total loss.

[0086] Specifically, based on the above total loss function, the sequence labeling loss, boundary detection loss and relationship prediction loss are integrated to fully reflect the performance of the model in terms of label prediction accuracy, entity boundary positioning accuracy and semantic relationship prediction accuracy. When the total loss is large, it indicates that the model has a large deviation in performance in these aspects.

[0087] Moreover, the total loss can dynamically balance the priorities of multiple tasks such as sequence labeling, boundary detection, and relationship prediction. For example, in nested entities (such as "acute exacerbation of chronic obstructive pulmonary disease"), the power parameters (η, θ, μ) can amplify the weight of semantic relationships (μ>1), while the gating parameters (δ) can suppress the interference of boundary noise (θ<1), and the back propagation algorithm uniformly adjusts the parameters of BERT, CNN, Transformer, and CRF layers according to the gradient of the total loss, so that the model gradually converges to a state of coordinated optimization of labels, boundaries, and semantic relationships, and finally achieves accurate recognition and robust generalization of complex entities (such as non-continuous, multi-level structures).

[0088] For label sequence optimization, if the sequence annotation loss accounts for a large proportion, it means that the predicted label sequence is significantly different from the actual label sequence. At this time, it may be necessary to adjust the parameters of the conditional random field layer, such as the transfer probability matrix, so that the generation of the label sequence is more in line with the actual situation; if the boundary detection loss is prominent, it is necessary to re-examine the long entity structure model and the feature extraction part, such as checking whether the long-distance dependency relationship determined by the Transformer model is accurate, and whether the local character combination features extracted by the convolutional neural network are complete. By adjusting the parameters or network structure of these parts, the accuracy of entity boundary positioning can be improved; if the relationship prediction loss is large, it may be necessary to optimize the process of generating a long entity structure model based on long-distance dependency relationships, or to fine-tune the feature extraction method of the BERT model to enhance the ability to capture semantic relationships.

[0089] For the overall behavior of the model, depending on the total loss, it is necessary to adjust hyperparameters such as the learning rate and number of iterations, or make appropriate adjustments to the overall architecture of the model, such as adjusting the connection method between different layers, so that the model can be continuously optimized in the subsequent training and prediction process, and ultimately improve the overall performance of named entity recognition, and more accurately determine the named entities corresponding to the text data.

[0090] Step 211, determining the named entity corresponding to the text data according to the adjusted optimized tag sequence.

[0091] Specifically, after the adjustment process in the above steps, an adjusted optimized label sequence is obtained. Each label in the adjusted optimized label sequence is checked and analyzed in turn through the CRF layer. For example, in the BIO annotation system commonly used in named entity recognition, if a label starting with "B" (Begin, i.e., the beginning of the entity) is encountered, such as "B-PERSON", it means that a person name entity is marked here, and then the subsequent labels will be checked. If it is followed by continuous "I" (Inside, i.e., inside the entity) labels, such as "I-PERSON", it means that the tokens corresponding to these positions are all parts of the same person name entity, and they will be collected in turn.

[0092] Then, continue scanning along the optimized label sequence according to this rule. Once you encounter the "O" (Other, i.e. non-entity) label or the next different type of entity label starting with "B", it is determined that the previously collected group of consecutive tokens with corresponding entity start and internal labels together constitute a complete named entity. For example, the group of tokens corresponding to the "B-PERSON" and "I-PERSON" mentioned above are combined to form a specific person's name.

[0093] In this way, by traversing the entire optimized label sequence from beginning to end, different types of named entities are continuously identified and extracted according to the label categories and arrangement rules, and finally all the named entities contained in the text data are determined, realizing accurate mapping and extraction from the optimized label sequence to specific named entities.

[0094] This embodiment uses the above method steps to first use the BERT model to perform feature extraction processing on the text data to be processed, and obtain multiple tokens and corresponding semantic feature vector sequences. This step can preliminarily parse the semantic information of the text and lay the foundation for subsequent processing. Then, convolution operations of different scales are performed through the convolutional neural network to obtain local character combination features, and they are fused with the semantic feature vector sequence to form a feature-enhanced semantic feature vector sequence. This process can integrate local and overall semantic details and mine semantic associations at different levels in the text.

[0095] Then, the Transformer model uses the attention mechanism to determine the long-distance dependencies of semantic feature vectors in the feature-enhanced semantic feature vector sequence. With its powerful capture ability, even for texts with non-continuous entities and composite entities such as "large, medium and small generators", it can span the intervals between words and accurately grasp the semantic associations between the various parts such as "large generators" and "medium-sized generators" as well as their positional relationships in the text.

[0096] Generating a long entity structure model based on such long-distance dependencies can construct an entity structure that fits the actual situation based on the determined complex semantics and positional relationships, and clearly present the internal structure and mutual connections of non-continuous and complex entities.

[0097] Finally, the conditional random field layer is used to optimize the label sequence of the relationship-enhanced semantic feature vector sequence and determine the named entities corresponding to the text data. Taking advantage of its label sequence optimization, accurate labels are rigorously and completely labeled for non-continuous entities and composite entities, thereby achieving accurate, rigorous, and complete labeling of such entities and solving the corresponding labeling problems. The performance of the NER task is improved, making it impossible to further improve the accuracy and completeness of the model output results.

[0098] The embodiment of the present invention also provides a relationship-enhanced named entity recognition method, which specifically includes the following steps: Step 301, performing feature extraction processing on the text data to be processed through the BERT model to obtain multiple tokens and a semantic feature vector sequence composed of the semantic feature vectors of the multiple tokens.

[0099] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0100] Step 302, sliding on the semantic feature vector sequence through convolution kernels of different widths in the convolutional neural network to generate feature maps of multiple scales.

[0101] The feature map is used to represent the semantic information of the semantic feature vector sequence at different scales.

[0102] Specifically, the convolution kernel can be regarded as a scanner with a specific "field of view". Convolution kernels of different widths represent different observation ranges and implement convolution operations of different scales. When these convolution kernels of different widths are applied to the sequence of semantic feature vectors, they will slide on this sequence at a certain step.

[0103] For example, a narrower convolution kernel covers relatively fewer vector elements in the semantic feature vector sequence each time, just like focusing on the vectors corresponding to several adjacent tokens. This can capture very close local semantic associations, reflect the semantic features within a small range, and generate a feature map at a certain scale. This feature map reflects the combination of semantic information in a small area.

[0104] A wider convolution kernel can cover more vector elements at one time and obtain semantic relationships with a longer span and a more macro level. For example, it can cover the semantic vectors of multiple tokens corresponding to a slightly longer phrase, and then generate another larger-scale feature map, showing a wider range of semantic structural characteristics.

[0105] As convolution kernels of different widths continue to slide over the entire semantic feature vector sequence, feature maps of multiple scales can be generated. These feature maps fully present the semantic information of the semantic feature vector sequence at different scales from different ranges and angles, providing a rich local feature foundation for subsequent more comprehensive and detailed understanding of text semantics and named entity recognition and other tasks.

[0106] Step 303: Activate the feature map through an activation function to obtain the local character combination feature.

[0107] Specifically, after generating feature maps of multiple scales, further processing is required to obtain the final local character combination features, and this is when the activation function is needed. The activation function is a nonlinear function that can bring nonlinear transformation capabilities to the feature map, allowing the model to learn more complex and abstract semantic representations.

[0108] For example, the commonly used ReLU activation function. When the element values ​​in the feature map are input into the activation function, the values ​​less than 0 will be set to 0, and the values ​​greater than 0 will be output as they are. Through such a transformation, the originally relatively flat and linear feature map information is given nonlinear characteristics, which can highlight more critical and distinctive semantic features.

[0109] After being processed by the activation function, the semantic information contained in the feature maps of different scales is further activated and refined, and some redundant or less important information is removed, thereby converting them into truly valuable forms that can reflect the local character combination features in the text. These activated features ultimately constitute the local character combination features we need, providing more expressive and targeted semantic materials for subsequent operations such as fusion with other features, helping to advance the entire named entity recognition process more accurately.

[0110] Step 304: fuse the local character combination feature with the semantic feature vector sequence to obtain a feature-enhanced semantic feature vector sequence.

[0111] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0112] Step 305: Through the Transformer model, an attention window is used to slide on the feature-enhanced semantic feature vector sequence to obtain multiple initially recognized entities.

[0113] First, the feature-enhanced semantic feature vector sequence is used as the input data of the Transformer model. The attention window in the model is like a flexible and focusing "observation frame", which will start to slide on the entire feature-enhanced semantic feature vector sequence according to certain rules. When sliding, the vectors in the attention window will perform deep interactive calculations through the attention mechanism. For example, for each local vector range covered by the window, the degree of correlation between the semantic feature vectors will be measured, and their connection in semantics and position will be analyzed to determine whether it is possible to form an entity.

[0114] For example, when processing the feature vector sequence corresponding to a text about a technology product introduction, when the attention window slides to certain areas, it is found that the tokens corresponding to the vectors have close semantic relationships such as modification, limitation or belonging. The vectors corresponding to tokens such as "smart" and "mobile phone" show correlations in the window that are consistent with the entity structure, so it is preliminarily determined that it may be an entity.

[0115] As the attention window continues to slide across the entire sequence, such analysis and judgment are performed on each local area in turn, and then multiple initially recognized entities that meet the corresponding conditions are screened out, laying the foundation for further determining long-distance dependencies and complete named entity recognition.

[0116] Step 306: determine the long-distance dependency relationship according to the entity position and semantic relationship of each of the initially recognized entities.

[0117] Specifically, from the perspective of entity position, the order of each initially identified entity in the entire text and their relative positions will be clarified. For example, in a text describing a commercial activity, the two initially identified entities "event organizer" and "event venue" are identified. By analyzing their positions in the text, we know that "event organizer" comes first and "event venue" comes later. This order of positions is a manifestation of positional relationship. From the perspective of semantic relationship, we will consider whether there is a semantic connection between the initially identified entities, whether it is a causal relationship, a parallel relationship, or a belonging relationship. For example, the two initially identified entities "enterprise" and "subsidiary products" have a semantic relationship of belonging, and "enterprise" plays a role in limiting the belonging of "subsidiary products".

[0118] By integrating such positional and semantic relationships between the initially recognized entities, we construct a context in which they are connected and influence each other, that is, we determine the long-distance dependencies. This enables the model to grasp the complex connections between different entities in the text, and provide key information support for subsequent generation of long entity structure models and more accurate named entity recognition tasks.

[0119] In some embodiments, step 306 specifically includes: adjusting the size of the attention window according to the entity position to obtain an adjusted attention window; using the adjusted attention window to slide on the feature-enhanced semantic feature vector sequence to determine the re-identified entity; determining the long-distance dependency relationship based on the entity position relationship and semantic relationship of the re-identified entity.

[0120] Specifically, once the initial entity position information is obtained, the size of the attention window will be adaptively adjusted based on this. For example, when an entity is identified to be in a key semantic core area in the text, and there may be other entity information closely related to it around it, in order to capture these related contents more comprehensively, the size of the attention window will be appropriately expanded to cover a wider range, so that more tokens around the entity and their corresponding semantic feature vectors can be included, without missing any potential semantic association clues. On the contrary, if the location of an entity is relatively independent and there is less information strongly associated with it around it, then the attention window will be correspondingly reduced to focus on the entity itself and its nearest neighboring area, and pay more attention to its local features. By dynamically and flexibly adjusting the size of the attention window according to the entity position, the adjusted attention window can better fit the actual situation of different entities in the text, and prepare for further mining of potential entities and their relationships.

[0121] After obtaining the adjusted attention window, it is allowed to slide in an orderly manner on the feature-enhanced semantic feature vector sequence. During the sliding process, each time the attention window covers an area, it will be deeply analyzed with the help of its internal attention mechanism. Since this window has been adjusted based on the entity position, it can capture information that conforms to the entity characteristics more specifically during the sliding process. For example, when the window covers a certain segment area, by considering the relationship between the semantic feature vectors therein, such as judging whether there is a reasonable semantic collocation, modification and limitation relationship between the tokens corresponding to these vectors, if it is found that the conditions for forming a new entity are met, such as the existence of a typical entity composition pattern such as "brand name + product name", it will be identified as a re-identified entity. As the adjusted attention window continues to slide on the entire feature-enhanced semantic feature vector sequence, more entities that may have been missed or not accurately identified before can be continuously excavated, enriching the recognition results of the entities, and providing more materials for the subsequent in-depth analysis of the relationship between entities.

[0122] After the re-identified entities are determined, the entity position relationship and semantic relationship between them should be comprehensively considered to clarify the long-distance dependency relationship. In terms of entity position relationship, the order and relative distance of each re-identified entity in the text will be carefully sorted out. For example, in a text describing an activity process, the two re-identified entities "activity registration stage" and "activity holding stage" can be analyzed by analyzing their positions. The former precedes the latter in time order. This positional order is an important manifestation of positional relationship. At the semantic relationship level, the focus is on analyzing whether there are semantic connections such as causality, belonging, and parallel between the re-identified entities. For example, the two re-identified entities "company" and "new products under the company" have a clear belonging semantic relationship. By integrating and analyzing these entity position relationships and semantic relationships, a network of interconnected and mutually influential networks between the re-identified entities is constructed, thereby determining the long-distance dependency relationship. This relationship can help the model more comprehensively and deeply grasp the intricate internal connections between different entities in the text, provide key basis for subsequent tasks such as generating long entity structure models, and further improve the accuracy and completeness of named entity recognition.

[0123] Step 307: Generate a long entity structure model based on the long-distance dependency relationship.

[0124] In some embodiments, this step may also include: first determining text content features of the text data; and then determining model weight parameters of the positional relationship and the semantic relationship based on the text content features.

[0125] The weight parameter is used to characterize the importance of the position relationship and the semantic relationship in the long entity structure model.

[0126] The content characteristics of texts cover many aspects, such as the genre and style of the text. For example, technical documents often have a rigorous structure and standardized expressions, and the position relationship may be relatively more important; while literary works focus more on the richness and flexibility of semantics, and semantic relationships are particularly critical. It also includes the fields involved in the text. For example, in medical texts, there are unique requirements for specific positional collocations and accurate semantic associations between professional terms.

[0127] Next, we need to determine the model weight parameters of positional relationships and semantic relationships in the long entity structure model based on these text content features. If the text content features show that the text focuses on positional information such as logical order and sequence, then when constructing the long entity structure model, a higher weight parameter will be given to the positional relationship, which means that the model will rely more on the positional relationship to sort out the relationship between the various parts of the long entity during the processing process. For example, the positional information such as the order in which the symptoms appear in the description of medical conditions will be considered. On the contrary, if the text content features reflect that semantic richness and relevance are more critical, such as the semantic relationship plays a leading role in shaping the image of things in literary descriptions, then a higher weight parameter will be assigned to the semantic relationship, allowing the model to grasp the relationship between the long entity and the surrounding content based more on semantic connections when constructing the long entity structure model.

[0128] In this embodiment, the weight parameters clearly characterize the importance of positional relationships and semantic relationships in the long entity structure model. They can guide the model to reasonably utilize the two aspects of positional and semantic information based on the actual characteristics of different texts, and more accurately construct a structural model that conforms to the text content and helps to accurately identify long entities, thereby improving the effect of the entire named entity recognition task.

[0129] Step 308: Determine a relationship-enhanced semantic feature vector sequence according to the long entity structure model.

[0130] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0131] Step 309 , performing label sequence optimization processing on the relationship-enhanced semantic feature vector sequence through a conditional random field layer to obtain an optimized label sequence.

[0132] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0133] Step 310: determining the sequence labeling loss, boundary detection loss, and relationship prediction loss of the optimized label sequence.

[0134] In this embodiment, for the sequence labeling loss, the optimized label sequence is compared with the corresponding true label sequence. Through a specific loss calculation method, such as the commonly used cross entropy loss, the inconsistency between the predicted label and the true label is statistically analyzed. The obtained numerical value represents the degree of difference in the overall labeling accuracy. The larger the numerical value, the worse the model's performance in accurately giving the corresponding label for each position.

[0135] Specifically, the sequence labeling loss can be determined by designing a function of the sequence labeling loss, and the functional expression of the sequence labeling loss may include:

[0136] Among them, L seq represents the sequence labeling loss; n is the number of vectors; , represents the label sequence of the hypothesis prediction, represents the predicted label vector; Y = (y 1 , y 2 ,…, y n ), represents the true label sequence, y i represents the true label vector; seq Represents a hyperparameter, which is used to adjust the weight of the penalty term for the average vector distance between the predicted label sequence and the true label sequence. Hyperparameters are parameters that need to be manually set before training machine learning and deep learning models and cannot be automatically learned through the model training process. The choice of their values ​​will affect the performance and training results of the model. For example, the learning rate, number of iterations, regularization coefficient, etc. are all hyperparameters.

[0137] Get the true label vector y i The process includes: First, the text data used for training needs to be manually annotated. The annotator will annotate each word or character in the text according to the specific named entity annotation specification to determine the named entity category to which it belongs. For example, in a news text, for the sentence "Fruit Company released a new mobile phone", "Fruit Company" will be annotated as a named entity of the "organization" category, "new mobile phone" may be annotated as a named entity of the "product" category, and other words may be annotated as non-named entity categories.

[0138] These annotation information will form a label sequence, that is, the true label sequence Y = (y 1 , y 2 ,…, y n ). This vector y i It is usually represented in the form of One-Hot Encoding. For example, if there are C named entity categories, then it is a length of y iA vector of which only one element is 1 and the rest are 0. The position of 1 indicates the named entity category to which the word or character belongs.

[0139] Before inputting the labeled data into the model, the data needs to be preprocessed. This includes converting the text data into a format suitable for model input, such as segmenting the text, converting it into word vectors, etc. At the same time, the real label sequence is processed accordingly to ensure that y i Able to match and calculate with the output of the model.

[0140] Get the predicted label vector The process includes: After a series of feature extraction, feature fusion and model calculation, the relation-enhanced named entity recognition model predicts the named entity category for each word or character in the input text. Specifically, the last layer of the model (usually the fully connected layer) outputs a probability distribution, indicating the probability that each word or character belongs to each named entity category.

[0141] For the i-th word or character in the text, the model outputs a probability vector of length C , where each element represents the probability that the word or character belongs to the corresponding named entity category.

[0142] Then, the argmax function is usually used to select the position of the element with the highest probability as the predicted named entity category, thereby obtaining the predicted label vector This constitutes the label sequence of the hypothesis prediction .

[0143] In the function Represents the predicted label vector and the true label vector y i Some custom distance metric between them (e.g., Euclidean distance, cosine distance, etc., with appropriate adjustments).

[0144] Item 2 It is a distance penalty term between the average vector of the predicted label sequence and the true label sequence, which aims to strengthen the overall consistency.

[0145] By minimizing Lseq, the label sequence predicted by the model is made as close as possible to the true label sequence. The second penalty term ensures that the overall distribution of the predicted label sequence is consistent with the distribution of the true label sequence.

[0146] In this embodiment, the determination of boundary detection loss focuses on checking the definition of named entity boundaries in the optimized label sequence, matching and comparing them with the actual accurate entity boundaries, and calculating the degree of deviation of the model in determining where the named entity starts and ends. If there are many entity boundary division errors, the boundary detection loss value will be higher.

[0147] Specifically, it can be determined by the function expression of the boundary detection loss, and the function of the boundary detection loss may include:

[0148]

[0149]

[0150] Among them, L boundary represents the boundary detection loss; , represents the boundary prediction result; B =(b 1 , b 2 ,…, b n ), indicating the real boundary situation; represents the information entropy of the prediction boundary, represents the i-th predicted boundary situation; H(B) represents the information entropy of the real boundary, b i represents the i-th real boundary case; μ boundary Represents a hyperparameter used to adjust the weight of the sum of the absolute differences between the predicted boundary and the true boundary in the overall loss.

[0151] Get b i The process includes: Before training the named entity recognition model, a large amount of text data needs to be manually annotated. The annotator will carefully read the text to determine the exact boundary position of each named entity in the text. For example, in the sentence "Beijing is the capital of China", "Beijing" is a named entity, and the annotator will mark the start and end positions of the entity "Beijing".

[0152] For each named entity in the entire text dataset, such boundary annotation is performed. These annotation information will be converted into a sequence, that is, the actual boundary situation B=(b 1 ,b 2 ,…,b n ), where b i Indicates whether the i-th position is the boundary of a named entity. Usually, if the i-th position is the boundary of a named entity, b i May have the value 1; if not a boundary, it has the value 0.

[0153] Get The process includes: During the training process, the relation-enhanced named entity recognition model will learn the characteristics of text data and the boundary patterns of named entities. After a series of operations such as BERT model extracting features, convolutional neural network obtaining local character combination features, and Transformer model determining long-distance dependencies, the model will predict the boundaries of named entities in the text based on the learned knowledge.

[0154] Specifically, the model calculates the probability that each position is a named entity boundary based on the feature information of the input text. For example, for the i-th position in the text, the model outputs a probability value, indicating the possibility that the position is a named entity boundary. Then, by setting a threshold (such as 0.5), when the calculated probability value is greater than the threshold, the model will Set to 1, indicating that the position is predicted to be the boundary of the named entity; when the probability value is less than or equal to the threshold, Set to 0, indicating that the location is not predicted to be the boundary of the named entity. This gives the boundary prediction result .

[0155] The second term in the function , is the sum of the absolute differences between the predicted boundary and the true boundary. These two items are combined to measure the accuracy of boundary detection.

[0156] By minimizing L boundary , so that the boundary predicted by the model learning is as consistent as possible with the true boundary.

[0157] In this embodiment, for the relationship prediction loss, the semantic relationship between tokens in the potential named entities reflected by the optimized label sequence can be analyzed and compared with the semantic relationship that these tokens should have in the actual text. The relationship prediction loss is obtained by quantifying the difference between the two. For example, when the parts within the entity should be a modifying relationship but are mistakenly judged as other relationships, this loss value will increase.

[0158] This embodiment may define the relationship prediction loss in a manner based on matrix decomposition differences.

[0159] definition It is the singular value decomposition form of the predicted relationship matrix and the true relationship matrix. U, X and V, Y are orthogonal matrices; Σ, Ω are diagonal matrices.

[0160] Assume that the relationship prediction result is expressed as ; is the predicted relationship matrix, m is the number of rows in the matrix, and k is the number of columns in the matrix; The element in the i-th row and j-th column of the predicted relationship matrix represents the quantified value of the semantic relationship between the token contained in the predicted i-th potential named entity and the token contained in the j-th potential named entity. This quantified value of the semantic relationship can be coded in a discrete manner, such as assigning different integer codes to different relationship types, such as "management" relationship coded as 1, "collaboration" relationship coded as 2, etc.

[0161] The true relationship is expressed as , R is the true relationship prediction matrix; r ij It is the element in the i-th row and j-th column of the true relationship matrix R. It represents the quantitative value of the specific semantic relationship between the token contained in the i-th potential named entity and the token contained in the j-th potential named entity in the true relationship, reflecting the actual degree of semantic association.

[0162] The functional expression of the relationship prediction loss can be designed as follows:

[0163] Among them, L relation represents the relationship prediction loss; ν relation Represents a hyperparameter used to adjust the influence of the difference between the predicted relationship matrix and the true relationship matrix in the overall loss.

[0164] Specifically, U is the prediction relationship matrix The orthogonal matrix after singular value decomposition is: , Represents the prediction relationship matrix The element in the i-th row and j-th column of , where m is the number of rows and k is the number of columns of the matrix; X is the orthogonal matrix after the singular value decomposition of the real relationship matrix R, , r ij is the element in the i-th row and j-th column of the true relationship matrix R.

[0165] Σ is the prediction relationship matrix The diagonal matrix after singular value decomposition.

[0166] Ω is the diagonal matrix after the singular value decomposition of the real relationship matrix R.

[0167] V is the prediction relationship matrix The orthogonal matrix after singular value decomposition.

[0168] Y is the orthogonal matrix after the singular value decomposition of the true relationship matrix R.

[0169] Represents the prediction relationship matrix The singular value decomposition reconstruction form of Represents the singular value decomposition reconstructed form of the true relationship matrix R.

[0170] This embodiment measures the accuracy of relationship prediction by comparing the differences between the various parts after singular value decomposition and the differences in the reconstructed matrix.

[0171] By minimizing L relation , so that the relationship matrix predicted by the model learning is as close as possible to the true relationship matrix.

[0172] Step 311, determining the total loss according to the sequence labeling loss, the boundary detection loss, the relationship prediction loss and the total loss function.

[0173] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0174] Step 312, adjust the label sequence optimization processing and the overall behavior performance of the model according to the total loss.

[0175] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0176] Step 313: determine the named entity corresponding to the text data according to the adjusted optimized tag sequence.

[0177] For this step, please refer to the description in the previous embodiment and will not be repeated here.

[0178] A relationship-enhanced named entity recognition method proposed in this embodiment first performs feature extraction on text data through the BERT model to obtain multiple tokens and semantic feature vector sequences. Then, a convolutional neural network is used to perform convolution operations of different scales on the semantic feature vector sequence. Specifically, feature maps of multiple scales are generated by sliding convolution kernels of different widths on the sequence. These feature maps represent the semantic information at different scales, and then are processed by activation functions to obtain local character combination features. Then, the local character combination features are fused with the semantic feature vector sequence to obtain a feature-enhanced semantic feature vector sequence.

[0179] The Transformer model plays an important role in this process. On the one hand, the attention window is used to slide on the feature-enhanced semantic feature vector sequence to obtain the initial recognition entity, and then the long-distance dependency is determined based on the entity position and semantic relationship of the initial recognition entity. If further precision is required, the attention window size can be adjusted according to the entity position to obtain an adjusted attention window, which is used to slide on the sequence to determine the re-recognized entity, and the long-distance dependency is determined again based on the position relationship and semantic relationship of the re-recognized entity. On the other hand, the position query matrix of the position perception layer is used to determine the position of the token, and the position key matrix determines the relative position relationship between tokens.

[0180] When generating a long entity structure model based on long-distance dependencies, the text content features of the text data are also determined, and the weight parameters of the positional relationship and semantic relationship in the model are determined based on this feature. After that, the label sequence optimization processing of the relationship-enhanced semantic feature vector sequence is performed through the conditional random field layer to obtain an optimized label sequence. In this process, the sequence annotation loss, boundary detection loss, and relationship prediction loss are also determined to measure the difference between the predicted label sequence and the true label sequence, the deviation of the accuracy of the named entity boundary positioning, and the difference between the potential named entity semantic relationship prediction and the actual relationship. Finally, the label sequence optimization processing and the overall behavior performance of the model are adjusted according to these losses to determine the named entity corresponding to the text data. This method comprehensively utilizes a variety of technical means to comprehensively and accurately realize named entity recognition, improving the accuracy and reliability of recognition.

[0181] The method provided in the above embodiment can be executed by an electronic device. The following describes the electronic device in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , is a schematic diagram of a physical device structure of an electronic device in an embodiment of the present invention.

[0182] It should be noted that Figure 3 The structure of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0183] like Figure 3 As shown, the electronic device includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 to the random access memory (RAM) 403, such as executing the method described in the above embodiment. In RAM 403, various programs and data required for system operation are also stored. CPU 401, ROM 402 and RAM 403 are connected to each other through bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0184] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a liquid crystal display (LCD) and an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.

[0185] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, various functions defined in the present invention are performed.

[0186] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.

[0187] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings.

[0188] Specifically, the electronic device of this embodiment includes a processor and a memory, the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and one or more processors call the computer instructions to enable the electronic device to execute the method provided by the above embodiment.

[0189] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of the electronic device, the electronic device implements the method provided in the above embodiment.

[0190] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

[0191] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.

[0192] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.

Claims

1. A relation-enhanced named entity recognition method, characterized in that: include: Performing feature extraction processing on the text data to be processed by the BERT model to obtain multiple tokens and a semantic feature vector sequence composed of the semantic feature vectors of the multiple tokens; Performing convolution operations of different scales on the semantic feature vector sequence through a convolutional neural network to obtain local character combination features; The local character combination feature is merged with the semantic feature vector sequence to obtain a feature-enhanced semantic feature vector sequence; Determine the long-distance dependency relationship of the semantic feature vectors in the feature-enhanced semantic feature vector sequence by using the attention mechanism through the Transformer model, wherein the long-distance dependency relationship is used to characterize the positional relationship and semantic relationship between the tokens; generating a long entity structure model based on the long-distance dependency; Determine a relationship-enhanced semantic feature vector sequence according to the long entity structure model; Through the conditional random field layer, performing label sequence optimization processing on the relationship-enhanced semantic feature vector sequence to obtain an optimized label sequence; Determine the sequence labeling loss, boundary detection loss and relationship prediction loss of the optimized label sequence; the sequence labeling loss is used to measure the difference between the predicted label sequence output by the overall model and the real label sequence; the boundary detection loss is used to evaluate the deviation degree of the accuracy of the boundary positioning of the named entity from the conditional random field layer to the overall model; the relationship prediction loss is used to quantify the difference between the semantic relationship prediction and the actual relationship between the tokens contained in the potential named entity from the conditional random field layer to the overall model; Determine the total loss according to the sequence labeling loss, the boundary detection loss, the relationship prediction loss and the total loss function; The expression of the total loss function includes: ; Among them, L total represents the total loss, L seq represents the sequence labeling loss, L boundary represents the boundary detection loss, L relation represents the relationship prediction loss, α, β, γ are weight coefficients, λ1, λ2, λ3 are normalization coefficients, δ is a nonlinear gating parameter, η, θ, μ are nonlinear strength control parameters; Adjusting the label sequence optimization processing and the overall behavior performance of the model according to the total loss; The named entity corresponding to the text data is determined according to the adjusted optimized tag sequence.

2. The method according to claim 1, characterized in that: The step of performing convolution operations of different scales on the semantic feature vector sequence through a convolutional neural network to obtain local character combination features includes: Sliding on the semantic feature vector sequence by using convolution kernels of different widths in the convolutional neural network to generate feature maps of multiple scales, wherein the feature maps are used to represent semantic information of the semantic feature vector sequence at different scales; The feature map is activated by an activation function to obtain the local character combination feature.

3. The method according to claim 1, characterized in that The step of determining the long-distance dependency relationship of the semantic feature vectors in the feature-enhanced semantic feature vector sequence by using the attention mechanism through the Transformer model includes: By using the Transformer model, an attention window is used to slide on the feature-enhanced semantic feature vector sequence to obtain a plurality of initially recognized entities; The long-distance dependency relationship is determined according to the entity position and semantic relationship of each of the initially recognized entities.

4. The method according to claim 3, characterized in that The step of determining the long-distance dependency relationship according to the entity position and semantic relationship of each of the initially recognized entities comprises: Adjust the size of the attention window according to the entity position to obtain an adjusted attention window; Sliding the adjusted attention window on the feature-enhanced semantic feature vector sequence to determine a re-identified entity; The long-distance dependency relationship is determined according to the entity position relationship and the semantic relationship of the re-identified entity.

5. The method according to claim 1, characterized in that The functional expression of the sequence labeling loss includes: Among them, L seq represents the sequence labeling loss; n is the number of vectors; , represents the label sequence of the hypothesis prediction, Represents the predicted label vector; Y = (y1, y2,…, y n ), represents the true label sequence, y i represents the true label vector; seq Represents a hyperparameter.

6. The method according to claim 1, characterized in that The functional expression of the boundary detection loss includes: Among them, L boundary represents the boundary detection loss; , represents the boundary prediction result; B = (b1, b2,…, b n ), indicating the real boundary situation; represents the information entropy of the prediction boundary, represents the i-th predicted boundary situation; H(B) represents the information entropy of the real boundary, bi represents the i-th real boundary situation; μ boundary Represents a hyperparameter.

7. The method according to claim 1, characterized in that The functional expression of the relationship prediction loss includes: Among them, L relation represents the relationship prediction loss; ν relation represents a hyperparameter; U is the prediction relationship matrix The orthogonal matrix after singular value decomposition is: , Represents the prediction relationship matrix The element in the i-th row and j-th column of , where m is the number of rows and k is the number of columns of the matrix; X is the orthogonal matrix after the singular value decomposition of the real relationship matrix R, , r ij is the element in the i-th row and j-th column of the true relationship matrix R; Σ is the prediction relationship matrix The diagonal matrix after singular value decomposition; Ω is the diagonal matrix after singular value decomposition of the real relationship matrix R; V is the prediction relationship matrix The orthogonal matrix after singular value decomposition; Y is the orthogonal matrix after the singular value decomposition of the real relationship matrix R; Represents the prediction relationship matrix The singular value decomposition reconstruction form of ; Represents the singular value decomposition reconstructed form of the true relationship matrix R.

8. An electronic device, characterized in that: including one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program codes, wherein the computer program codes include computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed on an electronic device, the electronic device is caused to execute the method as claimed in any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is executed on an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Manual intervention word segmentation method for improving ASR recognition effect

    CN120833785A

  • Earthquake emergency plan named entity identification method and system

    CN121168456A