Two-stage optimized wireless network optimization field long entity recognition method and system

Through a two-stage optimization strategy and semantic similarity evaluation, the problems of insufficient model generalization ability and poor recognition effect in long entity recognition in the field of wireless network optimization are solved, and more efficient long entity recognition and evaluation are achieved.

WO2025201231A1PCT designated stage Publication Date: 2025-10-02BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
PCT/CN2025/084310
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-24
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the long entity recognition task in the field of wireless network optimization, existing technologies have problems such as insufficient model generalization ability and fuzzy entity boundaries leading to poor recognition effect. In particular, model tuning is difficult in small sample scenarios, and traditional evaluation indicators cannot accurately evaluate the effect of long entity extraction.

Method used

A two-stage optimization strategy is adopted. First, domain knowledge is learned through the entity type prediction task, using the pre-trained model TelBert. Then, entity-related semantic information is introduced into the machine reading comprehension framework, and a dual-pointer network is used to decode the entity. The recognition effect is evaluated by combining the semantic similarity evaluation indicator TelSimilar.

Benefits of technology

It improves the accuracy and generalization ability of long entity recognition in the field of wireless network optimization, alleviates the difficulty of model tuning in small sample scenarios, and provides a more comprehensive recognition evaluation method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084310_02102025_PF_FP_ABST
    Figure CN2025084310_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of wireless network optimization operations and maintenance, and provides a two-stage optimized wireless network optimization field long entity recognition method and system. The method comprises: using a pretrained long entity recognition model to process acquired text content to be recognized to obtain a long entity recognition result; by means of a first-stage predecessor task, acquiring a pretrained model TelBert having domain knowledge; and in a second stage, introducing semantic information related to an entity to obtain a machine reading comprehension framework-based long entity recognition model, and decoding the entity by means of a dual-pointer network. According to the present invention, knowledge in a specific field is learned by adding an entity type prediction task, the text representation learning capability of a base model is enhanced, and the difficulty of model tuning in a few-shot scenario is alleviated; the entity recognition model is improved to obtain an MRC-LER model suitable for document-level long entity recognition; and a semantic similarity-based evaluation index is proposed, and the effective extraction rate of entity key information is reasonably evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Two-stage optimized long entity recognition method and system for wireless network optimization Technical Field

[0001] The present invention relates to the technical field of wireless network optimization and operation and maintenance, and in particular to a two-stage optimized long entity recognition method and system in the wireless network optimization field. Background Art

[0002] With the rapid development of society and the continuous advancement of communication network technology, communication service demands are becoming increasingly massive and concurrent. The overall planning of wireless network optimization and operation and maintenance systems is evolving towards diversification, intelligence, and high-speed development. This has accumulated a vast amount of case text data, providing operation and maintenance personnel with valuable experience in case analysis and resolution. However, the sheer volume of case documents and the complex domain expertise in the wireless network optimization field pose significant challenges to case reading and searching. Consequently, the need to build a knowledge graph for wireless network optimization cases is growing.

[0003] When constructing a domain case knowledge graph, entity recognition is an important foundation for building a high-quality knowledge graph. Entity recognition aims to extract predefined entities from unstructured text, including entity location and classification. Case texts in the wireless network optimization field are typically long documents, where entity categories include problem phenomena, anomaly causes, solutions, processing results, device IDs, etc. Compared to general domain entities such as names of people and places, entities in this field are longer and are usually short sentences. In addition, the scale of manually annotated domain entity datasets is small, and the ability to learn the representation of domain document-level text using general training paradigms and deep learning methods is insufficient, resulting in poor entity recognition results.

[0004] Training paradigms refer to the different methods and strategies used to train deep neural networks. Each training paradigm has its own unique characteristics and applicable scenarios, including supervised learning, unsupervised learning, reinforcement learning, and "pre-training + fine-tuning." Supervised learning requires input and corresponding labeled training data. The model makes predictions by learning the relationship between input and label. In unsupervised learning, the training data is unlabeled, and the goal of model training is to discover hidden structures or patterns in the data. "Pre-training + fine-tuning" utilizes large-scale unsupervised datasets for pre-training, allowing the model to learn the general characteristics and knowledge of the data, and then fine-tunes it on a specific dataset for the downstream task.

[0005] Entity recognition is a crucial task in natural language processing (NLP) information extraction, aiming to extract structured information from unstructured text data. This task aims to identify and classify meaningful entities within text, such as names, places, and organizations. Entity recognition centers on two main tasks: identifying entity boundaries, which determine the start and end locations of entities within a text; and classifying entities, which categorizes identified entities into predefined categories. This provides the necessary foundation for many subsequent NLP tasks, including sentiment analysis, relationship extraction, knowledge graph construction, and question answering.

[0006] Text classification involves assigning a category to the content of a text description based on predefined topic categories. Text classification tasks are generally divided into two categories: single-label classification and multi-label classification. Single-label classification involves assigning a single output category to an input text, while multi-label classification involves assigning multiple output categories. In single-label classification, if there are only two possible categories, it is called binary text classification. If there are more than two possible categories, it is called multi-label text classification.

[0007] Currently, there are three main approaches to model training: fully supervised learning based on non-neural networks (matching and utilizing features to complete NLP tasks through manually set rules and statistics), fully supervised learning based on neural networks (designing appropriate neural network architectures for training on large amounts of labeled data), and "pre-training + fine-tuning" (pre-training on large-scale unsupervised datasets, followed by fine-tuning on datasets specific to downstream tasks). Models trained under the "pre-training + fine-tuning" paradigm have achieved tremendous success in a wide range of tasks in natural language processing, becoming one of the mainstream technologies. Pre-training on large-scale unsupervised datasets allows the model to learn the grammatical and semantic features inherent in general text and can be applied to various fields and tasks. Fine-tuning on datasets specific to downstream tasks makes the model more adaptable to downstream tasks and possesses strong generalization capabilities. The "pre-training + fine-tuning" paradigm has achieved excellent results in solving a wide range of NLP tasks, including NER. However, in highly specialized areas such as wireless network optimization, when the amount of domain data is insufficient to support model fine-tuning, using general-purpose pre-trained language models can easily lead to overfitting and poor model performance. Therefore, a certain amount of high-quality domain-labeled data is required for "fine-tuning." Manual labeling is time-consuming and labor-intensive, and if the data annotation quality is not high, model tuning is difficult.

[0008] Early entity recognition technologies relied on rule-based and dictionary-based approaches, relying heavily on large amounts of hand-crafted features and domain knowledge, resulting in weak generalization capabilities. Machine learning-based methods such as support vector machines (SVMs) and hand-coded multi-model models (HMMs) rely on large amounts of annotated data, poorly extracting text semantics, and exhibit low accuracy. With the widespread application of deep learning in natural language processing, deep learning-based methods have become a mainstream research direction in entity recognition. Currently, scholars have conducted extensive research on deep learning-based entity recognition, including sequence labeling, pointer decoding, and sequence generation, depending on the entity decoding method. Most existing entity recognition algorithms identify short entities in general domains, such as names of people, places, and organizations. Entities in the wireless network optimization domain are generally long phrases or short sentences, which are highly specialized. Relying solely on text sequence modeling makes it difficult to learn semantic information related to entity types, resulting in low accuracy in long entity recognition. Furthermore, sequence labeling-based methods are prone to chain breaks when decoding labels for long entities. Sequence generation-based methods, however, face efficiency issues when dealing with long texts, with a large number of candidate entities. Furthermore, the generated results are uncontrollable, making long entity recognition even more challenging.

[0009] Currently, entity recognition evaluation metrics include exact match evaluation and loose match evaluation, using accuracy, recall, and F1 (the harmonic mean of precision and recall) to assess the effectiveness of model extraction. In exact match evaluation, a prediction is considered correct only when the predicted entity type and entity boundary are completely accurate. In loose match evaluation, when the predicted entity type is accurate, the prediction is considered correct as long as the predicted entity boundary coincides with the target entity boundary. In practical applications, exact match evaluation is more widely used. Due to the non-standardization and diversity of case text descriptions in the wireless network optimization field, the boundaries of domain entities are ambiguous. Therefore, long entity recognition can be considered correct if the predicted entity type is accurate and the entity content does not affect semantic understanding. However, exact match evaluation metrics require strict identification of entity boundaries, while loose match evaluation metrics only require that the predicted entity boundary coincide with the target entity boundary. Neither of these metrics accurately evaluates the effectiveness of the model in extracting long entities. Summary of the Invention

[0010] The object of the present invention is to provide a two-stage optimized method and system for long entity recognition in wireless network optimization field, so as to solve at least one technical problem existing in the above background technology.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] In a first aspect, the present invention provides a two-stage optimized method for long entity recognition in wireless network optimization domain, comprising:

[0013] Get the text content to be recognized;

[0014] Process the acquired text content to be recognized using a pre-trained long entity recognition model to obtain a long entity recognition result; wherein, the training of the long entity recognition model includes: the first stage is the learning stage of the pre-task, that is, the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, that is, the long entity recognition task, introduce semantic information related to the entity in TelBert to obtain a long entity recognition model based on the machine reading comprehension framework, and decode the entity in the way of a double-pointer network to improve the performance of long entity recognition in the field of wireless network optimization.

[0015] Optionally, the entity type prediction task is modeled as a multi-classification task. First, semantically represent the entity entity i to construct [CLS]e1,e2,...e o [SEP], where [CLS] and [SEP] are special tokens used to mark the start and segmentation. Input the sequence into the pre-trained model BERT and output a [CLS] vector. Among them, k refers to the total number of entity categories. The [CLS] vector output passes through a semantic linear classifier to judge the type of entity i to obtain the pre-trained model TelBert.

[0016] Optionally, in the pre-tuning stage of the pre-task, continuously update the network parameters through the backpropagation algorithm. The goal is to minimize the error between the result predicted by the model and the true label, and continuously improve the generalization ability and prediction performance of the model; use the cross-entropy loss function to evaluate the error between the true label Label type and the model prediction value Pred type 之间的误差。

[0017] Optionally, the long entity recognition task is modeled as a machine reading comprehension task. The data set is constructed in the form of a (Query, Context, Answer) triple, where Query refers to a sentence related to the content of the text to be recognized Context, and Answer refers to the target entity corresponding to Query; for each entity type y ∈ Y, construct the question Q y =(q1,q2,...q m ), m represents the length of the question, and label the entity e start,end according to the entity category y, where e start,end is a sub-fragment in Context and start < end, and thus construct the triple (Q y , Context, e start,end ).

[0018] Optionally, first take the question Q yContext y By splicing the special delimiters [CLS] and [SEP], the constructed sequence is:

[0019] [CLS]q1,q2,...q m [SEP]c1,c2,...c n [SEP];

[0020] The TelBert model receives a sequence and outputs a representation matrix x represents the sequence length, d represents the vector dimension of the last layer of RoBERTa; since the scope of entity extraction is only in the context Context, it does not include the question Q y , so only the context representation matrix is ​​retained n represents the context length;

[0021] Optionally, a single domain extraction method is used to predict the start / end index with the highest probability given a query and context; the context representation matrix H output by the TelBert model is c , the model predicts the probability of each token being the starting position, the formula is as follows:

[0022] Logit start Indicates the probability of predicting the subscript as the starting position, Logit end Indicates the probability of predicting the end position, W start 、W end is the learning parameter;

[0023] Using the argmax function on Logit start and Logit end , thereby obtaining the starting position and ending position with the highest probability;

[0024] During training, the context Context corresponds to two sets of label sequences Y start 、Y end , Y start 、Y end The length of is consistent with the length of Context, indicating that each token in Context is the true label of the starting position and the ending position of the entity. Therefore, for the prediction of the starting position and the ending position, the loss function is expressed as follows:

[0025] Loss start =CE(Logit start ,Y start )

[0026] Loss end =CE(Logit end ,Y end )

[0027] Where CE represents the cross-entropy loss function.

[0028] In a second aspect, the present invention provides a two-stage optimized wireless network optimization domain long entity recognition system, comprising:

[0029] An acquisition module is used to obtain the text content to be recognized;

[0030] The processing module is used to use a pre-trained long entity recognition model to process the acquired text content to be recognized to obtain a long entity recognition result. The training of the long entity recognition model includes: a first stage is a learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

[0031] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the two-stage optimized long entity recognition method for wireless network optimization field as described in the first aspect is implemented.

[0032] In a fourth aspect, the present invention provides a computer device comprising a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the two-stage optimized wireless network optimization field long entity recognition method as described in the first aspect.

[0033] In a fifth aspect, the present invention provides an electronic device comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the two-stage optimized wireless network optimization field long entity recognition method as described in the first aspect.

[0034] The beneficial effects of the present invention are as follows: by adding entity type prediction tasks (pre-tasks) to learn specific domain knowledge, the ability of the base model to learn text representation is enhanced, and the difficulty of model tuning in small sample scenarios is alleviated; the entity recognition model is improved, and an MRC-LER model suitable for document-level long entity recognition is proposed; an evaluation index based on semantic similarity is proposed to reasonably evaluate the effective extraction rate of entity key information.

[0035] Additional aspects and advantages of the present invention will be set forth in part in the following description, will become apparent from the following description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] FIG1 is a flow chart of a two-stage optimized document-level long entity recognition method according to an embodiment of the present invention.

[0038] FIG2 is a structural diagram of a pre-task model according to an embodiment of the present invention.

[0039] FIG3 is a structural diagram of the MRC-LER model according to an embodiment of the present invention.

[0040] FIG4 is a pseudo code diagram of the similarity evaluation index according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The embodiments of the present invention provide anti-human β-amyloid monoclonal antibodies and their applications. Those skilled in the art can refer to the contents of this article and appropriately improve the process parameters to achieve them. It is particularly important to point out that all similar replacements and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications herein without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.

[0042] General pre-trained language models are difficult to "fine-tune" in scenarios with small domain samples. Existing technologies generally rely on large-scale labeled data sets to tune pre-trained models. However, for scenarios with small domain samples, the model is prone to overfitting. The technical solution of the present invention designs a two-stage optimization strategy. By adding entity type prediction tasks (pre-tasks) to learn domain knowledge, it not only uses the labeled entities of existing data sets, but also adds "professional terms" of the domain knowledge base, injecting high-quality domain semantic knowledge to a great extent, enhancing the ability of pre-trained models to learn domain text representations, and alleviating the difficulties of model tuning in small sample scenarios.

[0043] Existing technologies generally perform entity feature extraction and analysis solely through context sequences. While identifying entities of different categories through context sequence information, the model's ability to learn and mine implicit feature information is limited. The technical solution presented in this paper models the long entity recognition task as a machine reading comprehension task, introducing semantic knowledge related to entity categories to enhance the model's entity feature search capabilities. Combined with the decoding method of a dual-pointer network, this solution effectively addresses the "broken link" problem of long entities.

[0044] Classic evaluation indicators are not suitable for evaluating the performance of long entity extraction. The evaluation indicators of exact matching require strict identification of entity boundaries, and the evaluation indicators of loose matching only require that the predicted entity boundaries coincide with the target entity boundaries. However, due to the non-standardization and diversity of the text descriptions of cases in the field of wireless network optimization, the boundaries of domain entities are fuzzy, and both cannot accurately evaluate the effect of the model in extracting long entities. Without loss of generality, the present invention proposes an evaluation method based on semantic similarity, TelSimilar, for the task of extracting long entities in the domain. It is composed of similarities such as LCS, ED, and TF-IDF, and can reasonably evaluate the effective extraction rate of key information of long entities in the domain, so as to obtain more comprehensive model recognition evaluation results.

[0045] Example 1

[0046] In this embodiment 1, a two-stage optimized long entity recognition system in the wireless network optimization field is first provided, including: an acquisition module for acquiring text content to be recognized; a processing module for processing the acquired text content to be recognized using a pre-trained long entity recognition model to obtain a long entity recognition result; wherein, the training of the long entity recognition model includes: a first stage is a learning stage of the pre-task, namely, the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely, the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

[0047] In this Embodiment 1, by using the above system, a method for long entity recognition in the field of wireless network optimization with two-stage optimization is implemented, including: using an acquisition module to acquire the text content to be recognized; using a processing module to process the acquired text content to be recognized with a pre-trained long entity recognition model to obtain a long entity recognition result; wherein, the training of the long entity recognition model includes: the first stage is the learning stage of a pre-task, that is, an entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, that is, the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in the manner of a double-pointer network to improve the performance of long entity recognition in the field of wireless network optimization.

[0048] The entity type prediction task is modeled as a multi-classification task. First, the entity i is semantically represented and constructed as [CLS]e1,e2,...e o [SEP], where [CLS] and [SEP] are special tokens used to mark the start and segmentation. The sequence is input into the pre-trained model BERT, and a [CLS] vector is output. Among them, k refers to the total number of entity categories. The [CLS] vector output passes through a semantic linear classifier to judge the type of entity i and obtain the pre-trained model TelBert.

[0049] In the pre-tuning stage of the pre-task, the network parameters are continuously updated through the backpropagation algorithm. The goal is to minimize the error between the result predicted by the model and the true label, and continuously improve the generalization ability and prediction performance of the model; the cross-entropy loss function is used to evaluate the error between the true label Label type and the predicted value Pred type of the model. The long entity recognition task is modeled as a machine reading comprehension task. The data set is constructed in the form of a (Query, Context, Answer) triple, where Query refers to a sentence related to the content of the text Context to be recognized, and Answer refers to the target entity corresponding to Query; for each entity type y∈Y, a question Q y =(q1,q2,...q m ) is constructed, m represents the length of the question, and the entity e start,end is annotated according to the entity category y, where e start,end is a sub-fragment in Context and start<end, and thus a triple (Q y , Context, e start,end ) is constructed.

[0050] First, the question Q yContext y By splicing the special delimiters [CLS] and [SEP], the constructed sequence is:

[0051] [CLS]q1,q2,...q m [SEP]c1,c2,...c n [SEP];

[0052] The TelBert model receives a sequence and outputs a representation matrix x represents the sequence length, d represents the vector dimension of the last layer of RoBERTa; since the scope of entity extraction is only in the context Context, it does not include the question Q y , so only the context representation matrix is ​​retained n represents the context length;

[0053] Using the single domain extraction method, given a query and context, the start / end index with the highest probability is predicted; the context representation matrix H output by the TelBert model c , the model predicts the probability of each token being the starting position, the formula is as follows:

[0054] Logit start Indicates the probability of predicting the subscript as the starting position, Logit end Indicates the probability of predicting the end position, W start 、W end is the learning parameter;

[0055] Using the argmax function on Logit start and Logit end , thereby obtaining the starting position and ending position with the highest probability;

[0056] During training, the context Context corresponds to two sets of label sequences Y start 、Y end , Y start 、Y end The length of is consistent with the length of Context, indicating that each token in Context is the true label of the starting position and the ending position of the entity. Therefore, for the prediction of the starting position and the ending position, the loss function is expressed as follows:

[0057] Loss start =CE(Logit start ,Y start )

[0058] Loss end =CE(Logit end ,Y end )

[0059] Where CE represents the cross-entropy loss function.

[0060] Example 2

[0061] In this second embodiment, a two-stage optimized document-level long entity recognition method is proposed. First, in the first stage, a pre-trained model, TelBert, with domain knowledge is obtained by learning the pre-task. Then, in the second stage, semantic information related to the entity is introduced to propose a long entity recognition model based on a machine reading comprehension framework. Entities are decoded using a dual-pointer network approach, improving the performance of long entity recognition in the wireless network optimization domain. The specific implementation process is shown in Figure 1 and includes the following two steps: pre-task and long entity recognition.

[0062] Through the analysis of domain datasets and entities, this embodiment designs an entity type prediction task as a prerequisite task. Through the learning of entity type prediction, domain entity knowledge is integrated into the pre-trained model BERT, improving the model's ability to learn text representations in the domain and alleviating the difficulty of tuning the few-sample model.

[0063] The entity type prediction task can be modeled as a multi-classification task. The model structure is shown in Figure 2. First, the entity i Perform semantic representation and construct [CLS]e1,e2,...e o [SEP], where [CLS] and [SEP] are used to mark the special tokens of the start and segmentation. The sequence is input into the pre-trained model BERT and a [CLS] vector is output. Where k refers to the total number of entity categories, and the [CLS] vector output is passed through a semantic linear classifier to determine entity i type.

[0064] In the pre-task pre-tuning phase, the network parameters are continuously updated through the back-propagation algorithm, with the goal of minimizing the error between the model prediction results and the true labels, and continuously improving the generalization ability and prediction performance of the model. The cross entropy loss function is used to evaluate the true label. type and the model prediction value Pred type Therefore, for entity category prediction, the loss function is expressed as follows:

[0065] Loss type =CE(Pred type,Label type )

[0066] In this embodiment, the long entity recognition task is modeled as a machine reading comprehension task, and the dataset is constructed in the form of (Query, Context, Answer) triples, where Query refers to a sentence related to the content of the input text Context, and Answer refers to the target entity corresponding to Query. For each entity type y ∈ Y, construct the question Q y =(q1, q2,... q m ), m represents the length of the question, and label the entity e start,end , where e start,end is a sub-fragment in Context and start < end, and thus construct the triple (Q y , Context, e start,end ).

[0067] (1) Construction of the question Query

[0068] Considering that the wireless network optimization entity labels have clear semantic meanings, in order to avoid introducing too much entity-irrelevant information in question construction, such as "find the one with... in the text", which interferes with the semantic search of the model, this embodiment directly uses the entity label as the question Query.

[0069] (2) Model design

[0070] To enhance the entity semantic representation learning of the model under long texts, this embodiment proposes a long entity recognition model MRC-LER based on the machine reading comprehension architecture, which decodes long entities in a double-pointer manner by introducing semantic information related to entity types. In this embodiment, TelBert pre-tuned by a pre-task is used as the main architecture of the model to perform semantic representation on the input, and the model structure is shown in Figure 3.

[0071] First, splice the question Q y and the context Context y through the special delimiters [CLS] and [SEP] to construct a sequence as follows:

[0072] [CLS]q1, q2,... q m [SEP]c1, c2,... c n [SEP]

[0073] The TelBert model receives the sequence and outputs a representation matrix x represents the sequence length, and d represents the vector dimension of the last layer of RoBERTa. Since the range of entity extraction is only in the context Context and does not include the question Qy , so only the context representation matrix is ​​retained n represents the context length.

[0074] Two binary classifiers are used in the entity domain prediction layer, one predicting whether the token is the start position and the other predicting whether the token is the end position. Analysis of domain entity annotation data shows that the number of entities in each category in a single text is mostly 0 or 1. Therefore, this embodiment adopts a single domain extraction method to predict the start / end index with the highest probability given a query and context. The context representation matrix H output by the TelBert model is c , the model predicts the probability of each token being the starting position, the formula is as follows:

[0075] Logit start Indicates the probability of predicting the subscript as the starting position, Logit end Indicates the probability of predicting the end position, W start 、W end To learn the parameters.

[0076] Using the argmax function on Logit start and Logit end , thereby obtaining the starting position and ending position with the highest probability.

[0077] During training, the context Context corresponds to two sets of label sequences Y start 、Y end , Y start 、Y end The length of is consistent with the length of Context, indicating that each token in Context is the true label of the starting position and the ending position of the entity. Therefore, for the prediction of the starting position and the ending position, the loss function is expressed as follows:

[0078] Loss start =CE(Logit start ,Y start )

[0079] Loss end =CE(Logit end ,Y end )

[0080] Where CE represents the cross-entropy loss function.

[0081] Classic evaluation metrics are not suitable for evaluating the performance of long entity extraction in the wireless network optimization domain. Unlike short entity recognition in general domains, the goal of long entity recognition in the wireless network optimization domain is to effectively extract key information without compromising semantic understanding of the entity. Entities are long and have fuzzy boundaries. If the core content is extracted, the entity can be considered valid. Therefore, without loss of generality, this embodiment proposes a TelSimilar evaluation method for long entity extraction tasks in this domain. This method measures the effectiveness of the algorithm by calculating the semantic similarity between the predicted entity and the target entity.

[0082] This embodiment adopts semantic similarity algorithms such as LCS, ED, and TF-IDF. LCS refers to the longest common subsequence similarity algorithm, which is used to calculate the overlap between the target entity and the predicted entity. The longer the common substring, the higher the similarity between the two entities. ED refers to the edit distance similarity algorithm, which calculates the minimum number of edits required for two entities to reach consistency. Editing operations include insertion, deletion, and replacement. The smaller the number of edits, the greater the similarity between the two entities. The TF-IDF text similarity algorithm is used to evaluate the importance of specific words relative to the document. TF represents word frequency, and IDF represents inverse text frequency index. First, each predicted entity and target entity are segmented to obtain multiple segmentations for each entity, and then the TF value and IDF value of each segmentation are calculated. The importance of each segmentation relative to the entity is determined by TF-IDF calculation. The higher the TF-IDF value, the closer the semantics of the predicted entity and the target entity.

[0083] In order to measure the effective extraction rate of key information of domain long entities, the TelSimilar calculation formula based on semantic similarity formed by similarity calculation combination features such as LCS, ED, and TF-IDF is as follows:

[0084] Among them, Similar value is Max(Similar LCS , Similar ED , Similar TF*IDF )

[0085] The pseudo code is shown in Figure 4.

[0086] Example 3

[0087] This embodiment 3 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the two-stage optimized wireless network optimization domain long entity recognition method is implemented. The method includes:

[0088] Get the text content to be recognized;

[0089] The acquired text content to be recognized is processed using a pre-trained long entity recognition model to obtain a long entity recognition result. The training of the long entity recognition model includes: the first stage is the learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

[0090] Example 4

[0091] This embodiment 4 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the two-stage optimized wireless network optimization domain long entity recognition method described above, the method comprising:

[0092] Get the text content to be recognized;

[0093] The acquired text content to be recognized is processed using a pre-trained long entity recognition model to obtain a long entity recognition result. The training of the long entity recognition model includes: the first stage is the learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

[0094] Example 5

[0095] This embodiment 5 provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the above-described two-stage optimized wireless network optimization domain long entity recognition method, the method including:

[0096] Get the text content to be recognized;

[0097] The acquired text content to be recognized is processed using a pre-trained long entity recognition model to obtain a long entity recognition result. The training of the long entity recognition model includes: the first stage is the learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

[0098] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0099] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0100] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0102] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solutions disclosed in the present invention without the need for creative work should be included in the scope of protection of the present invention.

Claims

1. A two-stage optimized method for long entity recognition in wireless network optimization domain, characterized in that: include: Get the text content to be recognized; The acquired text content to be recognized is processed using a pre-trained long entity recognition model to obtain a long entity recognition result. The training of the long entity recognition model includes: the first stage is the learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

2. The two-stage optimized wireless network optimization domain long entity recognition method according to claim 1 is characterized in that: The entity type prediction task is modeled as a multi-classification task. First, the entity i Perform semantic representation and construct [CLS]e1,e2,...e o [SEP], where [CLS] and [SEP] are used to mark the special tokens of the start and segmentation. The sequence is input into the pre-trained model BERT and a [CLS] vector is output. Among them, k refers to the total number of entity categories, and the [CLS] vector output passes through the semantic linear classifier to determine the entity i Type, get the pre-trained model TelBert.

3. The two-stage optimized wireless network optimization domain long entity recognition method according to claim 2 is characterized in that: In the pre-task pre-tuning phase, the network parameters are continuously updated through the back-propagation algorithm. The goal is to minimize the error between the model prediction results and the true labels, and continuously improve the generalization ability and prediction performance of the model. The cross entropy loss function is used to evaluate the true label. type and the model prediction value Pred type The error between .

4. The two-stage optimized wireless network optimization domain long entity recognition method according to claim 2 is characterized in that: The long entity recognition task is modeled as a machine reading comprehension task, and the dataset is constructed in the form of (Query, Context, Answer) triples, where Query refers to a sentence related to the content of the text Context to be recognized, and Answer refers to the target entity corresponding to Query; for each entity type y ∈ Y, construct the question Q y =(q1, q2,... q m ), m represents the length of the question, and annotate the entity e start,end according to the entity category y, where e start,end is a sub-fragment in Context and start < end, and thus construct the triple (Q y , Context, e start,end ).

5. The two-stage optimized wireless network optimization domain long entity recognition method according to claim 4 is characterized in that: First, the question Q y Context y Through the [CLS] and [SEP] special delimiters, the constructed sequence is: [CLS]q1,q2,...q m [SEP]c1,c2,...c n [SEP]; The TelBert model receives a sequence and outputs a representation matrix x represents the sequence length, d represents the vector dimension of the last layer of RoBERTa; since the scope of entity extraction is only in the context Context, it does not include the question Q y , so only the context representation matrix is ​​retained n represents the context length; 6. The two-stage optimized wireless network optimization domain long entity recognition method according to claim 5 is characterized in that: Using the single domain extraction method, given a query and context, the start / end index with the highest probability is predicted; the context representation matrix H output by the TelBert model c , the model predicts the probability of each token being the starting position, the formula is as follows: Logit start Indicates the probability of predicting the subscript as the starting position, Logit end Indicates the probability of predicting the end position, W start 、W end is the learning parameter; Using the argmax function on Logit start and Logit end , thereby obtaining the starting position and ending position with the highest probability; During training, the context Context corresponds to two sets of label sequences Y start 、Y end , Y start 、Y end The length of is consistent with the length of Context, indicating that each token in Context is the true label of the starting position and the ending position of the entity. Therefore, for the prediction of the starting position and the ending position, the loss function is expressed as follows: Loss start =CE(Logit start ,Y start ) Loss end =CE(Logit end ,Y end ) Where CE represents the cross-entropy loss function.

7. A two-stage optimized wireless network optimization domain long entity recognition system, characterized by: include: An acquisition module is used to obtain the text content to be recognized; The processing module is used to use a pre-trained long entity recognition model to process the acquired text content to be recognized to obtain a long entity recognition result. The training of the long entity recognition model includes: a first stage is a learning stage of the pre-task, namely the entity type prediction task, to obtain a pre-trained model TelBert with domain knowledge; in the second stage, namely the long entity recognition task, semantic information related to the entity is introduced into TelBert to obtain a long entity recognition model based on a machine reading comprehension framework, and the entity is decoded in a dual-pointer network manner to improve the performance of long entity recognition in the wireless network optimization field.

8. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the two-stage optimized wireless network optimization field long entity recognition method according to any one of claims 1 to 6 is implemented.

9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the two-stage optimized wireless network optimization field long entity recognition method according to any one of claims 1 to 6.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the two-stage optimized wireless network optimization field long entity recognition method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • BiLSTM-BiDAF named entity recognition method based on machine reading understanding

    CN114492441A

  • Multi-task language model-oriented meta-knowledge fine tuning method and platform

    WO2022088444A1

Cited By

  • Biological environment text named entity recognition method, medium, equipment and product

    CN121659944A

  • INP file analysis method

    CN121807316A

  • Multi-label text classification method based on semantic representation enhancement and dynamic weighted depolarization contrast learning and application thereof

    CN121858739A

  • Target operating system-oriented intention analysis and intelligent control method and system

    CN122222038A