Attention and MarkBERT fused earthquake prevention and disaster reduction entity identification method and system
By integrating attention with MarkBERT's deep learning model, the problem of low accuracy and recall of named entity recognition in the field of earthquake prevention and disaster reduction is solved, efficient entity recognition of earthquake prevention and disaster reduction text is achieved, and the model's ability to identify entity boundaries is enhanced.
Patent Information
- Application Number
- CN202510348123.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-29
AI Technical Summary
The existing named entity recognition technology has problems such as low recognition accuracy and recall, poor generalization capabilities of model, high computational complexity, difficulty in rapid deployment, and lack of effective modeling of word boundaries in the field of earthquake prevention and disaster reduction, especially in the field of long text or long-distance dependencies, which cannot effectively capture key context information.
The deep learning model that integrates attention and MarkBERT is adopted, including the MarkBERT encoding layer, timing analysis layer, dynamic context adjustment module and CRF decoding layer. By pre-processing, annotating and feature extraction of text data in the seismic field, and dynamically adjusting the weight of context information in combination with the self-attention mechanism, highlighting key domain entities, and weakening the impact of irrelevant information.
It improves the accuracy and effect of naming entity recognition, can extract entity information in earthquake prevention and disaster reduction text more accurately, solves the problem of insufficient feature information and low recognition efficiency, and enhances the model's ability to identify entity boundaries.
Smart Images

Figure CN120387449A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of entity recognition technology, and particularly to an earthquake prevention and mitigation entity recognition method and system that integrates attention and MarkBERT. Background Art
[0002] Named Entity Recognition (NER) is a basic task in natural language processing, mainly used to identify the categories and boundaries of entities in text. There are many deficiencies in the application of existing named entity recognition technologies in the field of earthquake prevention and mitigation.
[0003] Traditional rule - and statistics - learning - based methods often have difficulty accurately identifying boundaries when dealing with Chinese text, and the expression of text features is not sufficient, resulting in low accuracy and recall rates in recognition. At the same time, these methods highly rely on domain expert knowledge, and the generalization ability of the model is poor, making it difficult to adapt to other tasks or fields. In addition, although existing deep - learning - based models such as the combination of BiLSTM and CRF capture temporal features to a certain extent, when facing long texts or long - distance dependency relationships, the model often fails to effectively capture key context information, ignores key information, leading to a decline in recognition accuracy. At the same time, the model has a large number of training parameters, high computational complexity, and low training efficiency, making it difficult to be quickly deployed. Moreover, traditional models lack effective modeling of word boundaries, and it is easy to have problems of incomplete or incorrect entity recognition in complex earthquake prevention and mitigation texts, making it difficult to meet the requirements of high - precision entity recognition in the field of earthquake prevention and mitigation. Summary of the Invention
[0004] One or more embodiments of this specification provide an earthquake prevention and mitigation entity recognition method that integrates attention and MarkBERT, including:
[0005] S1. Collect the original text data in the earthquake field during a specific period, and pre - process the collected original text data;
[0006] S2. Label the pre - processed text data according to preset annotation objectives and rules, and extract features from the labeled data;
[0007] S3. Build a deep - learning model including a MarkBERT encoding layer, a temporal analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the labeled data;
[0008] S4. Use the test data set to evaluate the entity recognition performance of the trained model.
[0009] Furthermore, the original text data includes earthquake events, magnitudes, and earthquake spatio - temporal positions;
[0010] Collect the original text data in the field of earthquake in a specific period, and the preprocessing of the collected original text data specifically includes:
[0011] Review the original text data, modify the missing values, outliers and duplicate record data in the original text data, and convert the modified text data into a unified data format, including converting the date into a standard format and converting the unit of numerical data.
[0012] Furthermore, the annotation of the preprocessed text data according to the preset annotation objectives and rules specifically includes:
[0013] Entity type definition: Clearly define the entity types that need to be annotated, define different entity types in the named entity recognition task, including place names and organization names, and describe the definition and boundaries of each entity type;
[0014] Entity boundary annotation: Use character indexes or marker symbols to represent entity boundaries, and determine the start and end positions of the annotation;
[0015] Non-entity marking: Define the information that does not need to be annotated or should be excluded from the annotation scope, such as punctuation marks, stop words and label errors;
[0016] Consistency check: Based on the criteria and guidelines of consistency check, perform annotation consistency verification and correction during the annotation process.
[0017] Furthermore, the feature extraction of the annotated data includes the extraction of part-of-speech features, specifically:
[0018] Convert each word in the text into its part-of-speech tag, judge the entity boundary and context features according to the part-of-speech information, and determine the entity boundary through the context information of the words in the text.
[0019] Furthermore, the construction of the deep learning model including the MarkBERT encoding layer, the time series analysis layer, the dynamic context adjustment module and the CRF decoding layer specifically includes:
[0020] Use MarkBERT to encode the text to obtain a word vector tag sequence with word boundary information, and retain the semantic information and context relationship of the earthquake prevention and control text;
[0021] Pass the word vector sequence with word boundary information through BiLSTM, combine the context information, capture the time series features and dependencies in the text, and infer and annotate the earthquake prevention and mitigation entity sequence in the text;
[0022] Introduce a dynamic context adjustment module, dynamically adjust the weight of the context information according to the real-time semantic changes of the text, and highlight the key domain entities;
[0023] Responsible for modeling and decoding the tag sequence through CRF, and outputting the earthquake prevention and disaster reduction entity sequence.
[0024] Furthermore, the specific process of training the model using the labeled data includes:
[0025] Input the labeled training data into the model, calculate the loss function, and update the model's parameters using an optimization algorithm;
[0026] During the training process, select and adjust the hyperparameters of the model, including the learning rate and regularization parameters, and use cross-validation to select the best combination of hyperparameters.
[0027] Furthermore, the specific method for evaluating the entity recognition performance of the trained model using the test data set is:
[0028] Use the labeled test data to evaluate the trained model, calculate the performance metrics of the model in entity recognition. The performance metrics include accuracy, recall, and F1-score. The accuracy is the ratio of the number of correctly predicted entities by the model to the total number of entities; the recall is the ratio of the number of words correctly predicted as entities to the number of words of real entities; the F1-score is the harmonic mean that comprehensively considers precision and recall.
[0029] One or more embodiments of this specification provide an earthquake prevention and disaster reduction entity recognition system that integrates attention and MarkBERT, including:
[0030] Data processing module: used to collect the original text data in the earthquake field during a specific period and preprocess the collected original text data;
[0031] Feature extraction module: used to label the preprocessed text data according to preset annotation targets and rules, and extract features from the labeled data;
[0032] Model training module: used to build a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the labeled data;
[0033] Model testing module: used to evaluate the entity recognition performance of the trained model using the test data set.
[0034] One or more embodiments of this specification provide an electronic device, including:
[0035] A processor; and,
[0036] A memory arranged to store computer-executable instructions that, when executed, cause the processor to implement the steps of the above-mentioned earthquake prevention and disaster reduction entity recognition method that fuses attention and MarkBERT.
[0037] One or more embodiments of this specification provide a storage medium for storing computer-executable instructions that, when executed, implement the steps of the above-mentioned earthquake prevention and disaster reduction entity recognition method that fuses attention and MarkBERT.
[0038] By adopting the embodiments of the present invention, by placing the dynamic context adjustment module between BiLSTM and CRF and combining the accurate modeling of word boundaries by MarkBERT, the entity boundary recognition ability of the model is enhanced; the temporal features and dependencies of the text are extracted, combined with the self-attention mechanism, and the weights of the context information are dynamically adjusted in a timely manner according to the real-time semantic changes of the text, highlighting the key domain entities and weakening the influence of task-irrelevant or interfering information, so that the model can extract entity information more accurately when processing earthquake prevention and disaster reduction texts, effectively solving the problems of insufficient feature information and low recognition efficiency in the earthquake prevention and disaster reduction named entity recognition task, and improving the accuracy and effect of named entity recognition.
[0039] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 A flowchart of an earthquake prevention and disaster reduction entity recognition method that fuses attention and MarkBERT provided by one or more embodiments of this specification;
[0042] Figure 2 A model architecture diagram of an earthquake prevention and disaster reduction entity recognition method that fuses attention and MarkBERT provided by one or more embodiments of this specification;
[0043] Figure 3Schematic diagram of MarkBERT input for an earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT provided by one or more embodiments of this specification;
[0044] Figure 4 Schematic diagram of the LSTM unit structure for an earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT provided by one or more embodiments of this specification;
[0045] Figure 5 Schematic diagram of the composition of an earthquake prevention and disaster reduction entity recognition system integrating attention and MarkBERT provided by one or more embodiments of this specification;
[0046] Figure 6 Schematic diagram of the structure of an electronic device provided by one or more embodiments of this specification. Detailed implementation manners
[0047] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only some of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.
[0048] Method embodiments
[0049] According to an embodiment of the present invention, an earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT is provided. Figure 1 Flowchart of an earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT provided by one or more embodiments of this specification, as Figure 1 shown, the earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT according to an embodiment of the present invention specifically includes:
[0050] S1. Collect the original text data in the earthquake field during a specific period, and preprocess the collected original text data.
[0051] Collect the text data related to the earthquake field within a certain time range in a certain area. The original text data includes earthquake events, magnitudes, and earthquake spatio-temporal positions, etc., as the basis for model training and evaluation; specifically, the types of recognition data include the following:
[0052] Earthquake parameters: including earthquake magnitude, focal depth, epicenter position, earthquake occurrence time, etc.;
[0053] Seismic monitoring data: the location of seismic monitoring stations, monitoring data, sensor types, etc.;
[0054] Geological structure: including geological structure information such as plate boundaries, fault zones, seismic activity zones, etc.;
[0055] Geological properties: formation information, rock types, geological structure types, etc.;
[0056] Earthquake disasters: losses caused by earthquakes, affected areas, rescue information, etc.
[0057] Review the original text data, conduct a preliminary review of the original text data, check whether there are problems such as missing values, outliers, and duplicate records in the data, modify the missing values, outliers, and duplicate record data in the original text data, and convert and unify the data formats of the modified text data, such as converting dates to standard formats and converting the units of numerical data, etc., to ensure the consistency and usability of the data, including the following processing methods:
[0058] Handle missing values, check whether there are missing values in the data, which are usually represented by NaN or null values. According to the situation of missing values, you can choose to delete the rows containing missing values, fill in the missing values using interpolation methods (such as mean, median, or regression prediction), or perform other appropriate processing according to the characteristics of the data;
[0059] Handle outliers, through data analysis and visualization tools, discover possible outliers, which may deviate significantly from the data distribution. You can choose to delete the outliers, replace them with missing values and then process them, or perform appropriate processing according to specific domain knowledge. Check whether there are duplicate records in the data, that is, the situation where the rows are exactly the same. You can choose to delete the duplicate values to ensure that each record is unique;
[0060] Format conversion, convert the data into a suitable format to ensure the correct data type. For example, convert date-time data to date-time format and convert text data to numerical values, etc.;
[0061] Domain-specific processing, according to the characteristics of the problem and the domain to which the data belongs, perform corresponding data processing. For example, text data may need to be tokenized, stop-word processed, and stemmed, etc.;
[0062] Dataset division, divide the dataset into training set, validation set, and test set for model training, tuning, and evaluation.
[0063] S2. Annotate the preprocessed text data according to the preset annotation objectives and rules, and extract features from the annotated data.
[0064] First, clarify the annotation objectives and rules, and formulate a set of annotation specifications or guidelines to describe how to annotate specific types of information, specifically including: Definition of entity types: Clearly define the entity types to be annotated. In the named entity recognition task, define different entity types, including place names and organization names, and describe the definition and boundaries of each entity type; Annotation of entity boundaries: The entity boundaries can be represented using character indices or marking symbols to determine the start and end positions of the annotation; Non-entity markers: Define how to handle non-entity markers, and clarify the information that does not need to be annotated or should be excluded from the annotation scope, such as punctuation marks, stop words, and tagging errors; Consistency check: Based on the criteria and guidelines for consistency checking, perform annotation consistency verification and correction during the annotation process. In this embodiment, for earthquake parameter annotation: The earthquake magnitude should use the Richter magnitude, and the annotation format is "Mw X.X" (X.X is the magnitude value, for example, "Mw 7.0"); The focal depth annotation uses kilometers as the unit and is accurate to one decimal place; The epicenter location is annotated using latitude and longitude coordinates, in the format of "(latitude, longitude)", accurate to two decimal places; The occurrence time uses the ISO 8601 standard date and time format, such as "YYYY-MM-DDTHH:MM:SS". For earthquake monitoring data annotation: The monitoring station location annotation should include latitude and longitude coordinates and altitude, in the format of "(latitude, longitude, altitude)"; The monitoring data annotation needs to record the earthquake waveform data, including features such as amplitude, frequency, and duration. For earthquake disaster annotation: The affected area annotation should include the region name and the degree of damage, in the format of "region name:degree of damage"; The loss situation annotation needs to record the casualties, house collapses, economic losses, etc.
[0065] Feature extraction of the annotated data includes extracting part-of-speech features, specifically:
[0066] Convert each word in the text into its part-of-speech tag, such as noun, verb, adjective, etc. The annotator judges the entity boundaries and context features based on the part-of-speech information. Through the context information of the words in the text, such as the content and tags of the previous and next few words, the entity boundaries are determined. The context words can be used as features and input to the annotator so that they can obtain more comprehensive information, context relationships, etc. during annotation, and extract features from the annotated data for model training.
[0067] S3. Construct a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and use the annotated data to train the model.
[0068] Figure 2 The model architecture diagram of an earthquake prevention and mitigation entity recognition method that fuses attention and MarkBERT provided for one or more embodiments of this specification is as Figure 2As shown in the figure, the present invention targets the field of earthquake prevention and disaster reduction. The MarkBERT model structure is selected as the embedding layer, the BiLSTM is selected as the encoding layer, the dynamic context adjustment module is used as the feature fusion layer, and finally the CRF is used as the decoding layer to better process text data related to earthquake prevention and disaster reduction, such as earthquake early warning information, disaster rescue instructions, etc., and improve the performance of the model in this scenario.
[0069] The embedding layer is used to map the input discrete tokens (such as words or characters) to high-dimensional real-valued vectors. This helps the model capture the semantic information and context relationships in the input sequence. Especially in the scenario of earthquake prevention and disaster reduction, it is crucial to accurately understand the key information in the text. For example, for the professional terms and disaster descriptions in earthquake early warning information, the embedding layer can transform them into vector forms that the model can process. The present invention uses the MarkBERT pre-trained model to generate a sequence of character vectors with word boundary information. MarkBERT is a pre-trained language model that incorporates word boundary information. It uses the Transformer unit as the encoder, and the encoding unit obtains the weight values of the internal relationships of the sequence through multi-head attention. Compared with the traditional BERT model, MarkBERT inserts word boundary marker information between adjacent characters. By processing the word boundary information, the model can more accurately identify entities in earthquake prevention and disaster reduction texts such as earthquake magnitudes and names of affected areas, avoiding the problems of incomplete entity recognition or overfitting. This model takes characters as units to solve the out-of-vocabulary (OOV) problem caused by the vocabulary limit of BERT, which is of great significance for some newly emerged earthquake-related words or place names.
[0070] MarkBERT has made the following improvements compared with the BERT model: First, the text is tokenized, where PosTagggins is the part-of-speech of the corresponding word. In the scenario of earthquake prevention and disaster reduction, by annotating and processing the part-of-speech, the model can better understand the grammatical structure and semantic focus in the text. For example, it can distinguish between the verb "rescue" and the noun "rescue team", so as to more accurately extract information related to disaster prevention and reduction. The tokenization situation is shown in Table 1:
[0071] Table 1
[0072]
[0073] After tokenizing the text related to earthquake disasters, the part-of-speech information of each word is obtained, and a marker [S] consistent with the part-of-speech information of the word is added at the end of each word, which not only explicitly inputs the part-of-speech information but also can be used to distinguish word boundaries. The text with the added markers is input into MarkBERT, as Figure 3 shown.
[0074] MarkBERT has two advantages over other pre-trained language models:
[0075] (1) It is convenient to add word-level learning objectives to boundary markers, which complements traditional character- and sentence-level pre-training tasks. Especially when dealing with word-level important information such as earthquake-related terms, it can better learn their features;
[0076] (2) Richer semantics can be easily incorporated by replacing general markers with POS-tag-specific markers. For example, replacing the POS-tag markers of earthquake-related words with more earthquake-domain-specific markers enables the model to better understand their semantics.
[0077] The task of the encoding layer is to capture the context information of the input sequence for better context understanding. In the present invention, a bidirectional long short-term memory network is used to perform global feature extraction on the disaster prevention and mitigation character vector sequence output by the word embedding layer. The long short-term memory network LSTM is a type of recurrent neural network, which consists of a forget gate, an input gate, and an output gate. By forgetting information in the cell state and remembering new information, the information useful for subsequent moment calculations is passed on and the useless information is discarded. At the same time, the hidden layer state is output at each time. Figure 4 It is a schematic diagram of the BiLSTM network structure and the LSTM cell structure.
[0078] Calculate the forget gate to select the information to be forgotten. The input is the earthquake prevention and control data of the current module and the output h of the previous module t-1 :
[0079] f t = σ(W f *[h t-1 , x t +b f );
[0080] Among them, f t represents the output of the forget gate, W f represents the weight parameter, b f represents the bias parameter, h t-1 represents the hidden layer output of the previous moment, and x t represents the input of the current moment.
[0081] Calculate the input gate to select the information to be remembered. The input is the earthquake prevention and control data of the current module and the output h of the previous module t-1 :
[0082] i t = σ(W i *[h t-1 , x t +b i );
[0083]
[0084]
[0085] Among them, σ is the Sigmoid function, and W i represents the weight parameter in the neural network, and b i represents the bias parameter, and h t-1 represents the output of the hidden layer at the previous moment, and x t represents the input at the current moment, and W c and b c represent the weight and bias parameters, C t-1 and represent the module state at the previous moment and the candidate module state at the current moment, while C t represents the module state after memory update.
[0086] Calculate the output gate and the hidden layer state at the current moment, and select the output value:
[0087] O t = σ(W o * [h t-1 , x t + b o );
[0088] h t = O t * tanh(C t );
[0089] Among them, W o represents the weight parameter in the neural network, b o represents the bias parameter, O t represents the output result of the output gate at the current moment, and h t represents the hidden layer state at the current moment and is input to the subsequent module.
[0090] Since the LSTM model usually has a large number of parameters, which will cause the model to need to process a large amount of calculations in each training step, in this paper, singular value decomposition (SVD) is used to achieve sparsification to shorten the training time.
[0091] First, select the weight matrices of the input gate, forget gate, output gate, etc. to be sparsified, perform singular value decomposition on the selected weight matrices, and decompose them into the product of three matrices.
[0092] W = U * S * V T ;
[0093] Among them, U and V are two orthogonal matrices, S is a diagonal matrix. By retaining the larger singular values and setting the remaining singular values to zero, a low-rank approximation matrix is obtained to replace the original weight matrix.
[0094] From the structure of the LSTM, it can be found that the LSTM cannot fully utilize the text context information. Therefore, in this paper, a bidirectional long short-term memory network (BiLSTM) model is used to extract the key features for named entity recognition in earthquake prevention and control data. The BiLSTM model consists of a forward LSTM and a backward LSTM, which solves the problem that the LSTM cannot encode information from back to front, better captures bidirectional semantic dependency information, and obtains complete context information. The forward layer LSTM obtains the semantic features of the previous earthquake prevention and control data; the backward layer LSTM obtains the semantic features of the subsequent earthquake prevention and control data, and finally combines the two to obtain the semantic features of the earthquake prevention and control data.
[0095] The feature fusion layer ensures that the model can better process the important information in long texts, highlights the entities in the field of earthquake prevention and mitigation, and weakens the influence of task-irrelevant or interfering information. After the BiLSTM processing of the text is completed, the model has captured the temporal features and dependencies of the context. The attention mechanism then performs a preliminary calculation of the correlation between the words within the sequence. The input of the dynamic context adjustment module will include these two parts of information.
[0096] The module calculates the importance score of each word through a learnable context scoring mechanism based on its contribution to the current context. This mechanism depends on the local context window of the text and combines the temporal features output by the BiLSTM and the correlation information generated by the attention mechanism through a lightweight feed-forward neural network or attention scoring function to generate a set of context importance scores. The specific formula can be expressed as:
[0097] α t = softmax(W c [h t ; a t +b c );
[0098] Among them, h t is the output of the BiLSTM at time step t, α t is the output of the self-attention mechanism at time step t, W c and b c are trainable weight and bias terms, and α t is the context importance score at time step t.
[0099] According to the context importance scores of each word, the corresponding features are weighted and adjusted to enhance the lexical information that is crucial for tasks such as earthquake prevention and disaster reduction entity recognition, while weakening the information of unimportant or irrelevant words. For example, in earthquake disaster texts, key information such as "epicenter" and "rescue team" is enhanced, and irrelevant background information is weakened. This weighting operation can be achieved through the following formula:
[0100] h′ t =α t ×h t ;
[0101] where h′ t is the time series feature after dynamic context adjustment, which will be the input of the next CRF model.
[0102] Different sentences or text paragraphs may contain different semantic emergencies. Especially in earthquake prevention and control texts, key information often concentrates in certain specific segments. For example, an important shelter location may be suddenly mentioned in a text. Therefore, this module not only adjusts according to the context scores of the current word, but also allows the context window size to change dynamically to adapt to the characteristics of long or short sentences, thereby optimizing the processing of long texts. The adjustment of the window can be adaptively learned by the model according to historical context information, and the specific implementation method can be through the extension of the multi-head attention mechanism.
[0103] In earthquake prevention and control data, introducing a dynamic context adjustment module helps the model accurately identify key earthquake basics, earthquake prevention and disaster reduction, etc. information, solve many long-distance dependency relationships in the data, such as the relationship between earthquake magnitude and shelter locations. It helps the model effectively capture these dependencies and improve the model's ability to understand overall information. Such features dynamically adjusted by context are more targeted and can improve the accuracy and robustness of CRF for entity sequence decoding. Especially when processing texts in the field of earthquake prevention and control, it can more effectively distinguish key information and background information.
[0104] The decoding layer uses a conditional random field (CRF) to perform label prediction on the input sequence, that is, to classify named entities for each word or character. CRF can learn context information by considering the global information of the label sequence and adding constraints to the final prediction result, combining the global probability of the label sequence and the result of the output layer, and predicting the label sequence with the highest probability. Adjacent labels usually have strong dependencies in label tasks. To address this issue, this paper uses CRF to jointly decode the label information of a given input sentence. CRF is a sequence annotation model that takes into account the order and correlation between labels and has advantages in sequence annotation tasks. Therefore, given a prediction sequence y=(y1,y2,...,y n) Using sentence X, the CRF model obtains the best sequence of tags:
[0105]
[0106] where A and P are the transition score matrix and the output score matrix respectively, represents the transition score from tag i to tag i + 1, represents the output score yi of the i-th Chinese character i .
[0107] The Softmax function is used to normalize the scores on all possible tag paths to generate the conditional probability of each tag on path y:
[0108]
[0109] where, represents the true label value, Y x is all possible tag sequences. During training, to maximize P(y|x), during prediction, a set of sequences with the highest probability is output through the equation:
[0110]
[0111]
[0112] The NER model is trained using training data annotated with a large amount of text related to earthquake disasters. The annotated training data is input into the model, the loss function is calculated, and the parameters of the model are updated using an optimization algorithm. The training data is divided into small batches for training, and each batch contains multiple training examples. In each batch, the model performs forward propagation on the input data, calculates the loss, and calculates the gradient through backpropagation. Then, the optimization algorithm uses these gradients to update the model parameters to minimize the loss. Among them, the loss function can be the cross-entropy loss in multi-classification tasks or the conditional random field loss in sequence annotation tasks; during the training process, it is also necessary to select and adjust the hyperparameters of the model, including the learning rate and regularization parameters, and use the cross-validation method to select the best combination of hyperparameters.
[0113] S4. Use the test dataset to evaluate the entity recognition performance of the trained model.
[0114] Evaluate the trained model using the labeled test data and calculate the performance metrics of the model in identifying entities. First, prepare the test data set for evaluation. The test data set is data that the model has not seen during the training process and should include various texts related to earthquakes, such as earthquake monitoring reports, rescue instructions, disaster news, etc., to ensure that the evaluation results objectively reflect the true performance of the model. Perform the same preprocessing steps on the test data as on the training data, including data cleaning, transformation, and feature engineering, etc., to keep the preprocessing method consistent with that of the training data. Use the trained model to make predictions on the test data. The model gives corresponding output results according to the input data. In this embodiment, the output will be the entity labels of each word, such as earthquake magnitude, affected areas, relief supplies, etc.
[0115] Calculate the performance metrics of the model based on the prediction results of the model and the true labels of the test data. The performance metrics include accuracy, recall, and F1-score. The accuracy is the ratio of the number of earthquake prevention and disaster reduction-related entities correctly predicted by the model to the total number of entities; the recall is the ratio of the number of words correctly predicted as earthquake prevention and disaster reduction-related entities to the number of words of true earthquake prevention and disaster reduction-related entities; the F1-score is the harmonic mean that comprehensively considers precision and recall. The F1-score is more suitable for evaluating the performance of the model on imbalanced data, especially in the field of earthquake prevention and disaster reduction, where the data often has imbalance, such as relatively few but crucial earthquake early warning information.
[0116] Through the evaluation metrics, we can understand how well the model performs in the earthquake prevention and disaster reduction scenario. If the performance of the model meets the requirements, then it can be used in actual earthquake prevention and disaster reduction applications, such as auxiliary decision-making for earthquake early warning systems, rapid extraction of disaster information, etc. If the model performance is not good, it is necessary to further analyze the possible reasons and may need to perform model tuning or take other improvement measures, such as increasing earthquake prevention and disaster reduction-related training data, optimizing the model structure, etc.
[0117] To more intuitively understand the prediction results of the model, the output results of the model can be compared with the true labels and visualized, for example, through text marking or visualization of entity boundaries to highlight key information such as earthquake magnitude and affected area. If it is found that the model performance is still not ideal during the evaluation, model fine-tuning can be performed and re-evaluated until the predetermined performance requirements are met to ensure that the model can play an important role in earthquake prevention and disaster reduction work.
[0118] The beneficial effects of the present invention are as follows:
[0119] By adopting the embodiment of the present invention, by placing the dynamic context adjustment module between the BiLSTM and the CRF, and combining the accurate modeling of word boundaries by MarkBERT, the entity boundary recognition ability of the model is enhanced; the temporal features and dependencies of the text are extracted, and combined with the self-attention mechanism, the weight of the context information is dynamically adjusted in a timely manner according to the real-time semantic changes of the text, highlighting the key domain entities and weakening the influence of task-irrelevant or interfering information, so that the model can extract entity information more accurately when processing earthquake prevention and disaster reduction texts, effectively solving the problems of insufficient feature information and low recognition efficiency in the earthquake prevention and disaster reduction named entity recognition task, and improving the accuracy and effect of named entity recognition.
[0120] System embodiment
[0121] According to an embodiment of the present invention, there is provided an earthquake prevention and disaster reduction entity recognition system integrating attention and MarkBERT. Figure 5 As shown in the schematic diagram of the composition of an earthquake prevention and disaster reduction entity recognition system integrating attention and MarkBERT provided for one or more embodiments of this specification, Figure 5 As shown, the earthquake prevention and disaster reduction entity recognition system integrating attention and MarkBERT according to the embodiment of the present invention specifically includes:
[0122] Data processing module 50: used to collect the original text data in the earthquake field during a specific period, and preprocess the collected original text data;
[0123] Feature extraction module 52: used to label the preprocessed text data according to preset annotation targets and rules, and extract features from the labeled data;
[0124] Model training module 54: used to build a deep learning model including a MarkBERT encoding layer, a temporal analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the labeled data;
[0125] Model testing module 56: used to evaluate the entity recognition performance of the trained model using a test data set.
[0126] The embodiment of the present invention is a system embodiment corresponding to the above method embodiment. The specific operations of each module can be understood with reference to the description of the method embodiment, and will not be repeated here.
[0127] Device embodiment 1
[0128] The embodiment of the present invention provides an electronic device, as Figure 6 shown, including: a memory 60, a processor 62, and a computer program stored on the memory 60 and executable on the processor 62. When the computer program is executed by the processor 62, the following method steps are implemented:
[0129] S1. Collect the original text data in the field of seismology within a specific period, and preprocess the collected original text data;
[0130] S2. Annotate the preprocessed text data according to the preset annotation objectives and rules, and extract features from the annotated data;
[0131] S3. Construct a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the annotated data;
[0132] S4. Use the test data set to evaluate the entity recognition performance of the trained model.
[0133] Device Embodiment II
[0134] An embodiment of the present invention provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor 62, the following method steps are implemented:
[0135] S1. Collect the original text data in the field of seismology within a specific period, and preprocess the collected original text data;
[0136] S2. Annotate the preprocessed text data according to the preset annotation objectives and rules, and extract features from the annotated data;
[0137] S3. Construct a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the annotated data;
[0138] S4. Use the test data set to evaluate the entity recognition performance of the trained model.
[0139] The computer-readable storage medium described in this embodiment includes but is not limited to: ROM, RAM, magnetic disk, optical disc, etc.
[0140] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for earthquake prevention and disaster reduction entity recognition that integrates attention and MarkBERT, characterized in that Including: S1. Collect the original text data in the earthquake field within a specific period, and preprocess the collected original text data; S2. Annotate the preprocessed text data according to the preset annotation objectives and rules, and extract features from the annotated data; S3. Construct a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and use the annotated data to train the model; S4. Use the test data set to evaluate the entity recognition performance of the trained model.
2. The method according to claim 1, wherein The original text data includes earthquake events, magnitudes, and earthquake spatio-temporal locations; Collecting the original text data in the earthquake field within a specific period and preprocessing the collected original text data specifically includes: Review the original text data, modify the missing values, outliers, and duplicate record data in the original text data, and convert the modified text data into a unified data format, including converting the date into a standard format and converting the unit of numerical data.
3. The method according to claim 1, wherein The specific process of annotating the preprocessed text data according to the preset annotation objectives and rules includes: Entity type definition: Clearly define the entity types to be annotated, define different entity types in the named entity recognition task, including place names and organization names, and describe the definition and boundaries of each entity type; Entity boundary annotation: Use character indexes or marker symbols to represent entity boundaries, and determine the starting and ending positions of the annotation; Non-entity marking: Define the information that does not need to be annotated or should be excluded from the annotation scope, such as punctuation marks, stop words, and label errors; Consistency check: Based on the criteria and guidelines of the consistency check, perform annotation consistency verification and correction during the annotation process.
4. The method according to claim 3, characterized in that The feature extraction of the annotated data includes the extraction of part-of-speech features. Specifically: Convert each word in the text into its part-of-speech tag, judge the entity boundary and context features according to the part-of-speech information, and determine the entity boundary through the context information of the words in the text.
5. The method according to claim 1, wherein The specific process of constructing a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer includes: Use MarkBERT to encode the text to obtain a word vector marker sequence with word boundary information, and retain the semantic information and context relationship of the earthquake prevention and control text; Pass the word vector sequence with word boundary information through BiLSTM, combine the context information, capture the time series features and dependencies in the text, and infer and annotate the earthquake prevention and mitigation entity sequence in the text; Introduce a dynamic context adjustment module, dynamically adjust the weight of the context information according to the real-time semantic changes of the text, and highlight the key domain entities; Responsible for the modeling and decoding of the label sequence through CRF, and output the earthquake prevention and mitigation entity sequence.
6. The method according to claim 1, wherein The specific process of training the model with the annotated data includes: Input the annotated training data into the model, calculate the loss function, and use the optimization algorithm to update the parameters of the model; Select and adjust the hyperparameters of the model during the training process, including the learning rate and regularization parameters, and use the cross-validation method to select the best hyperparameter combination.
7. The method according to claim 1, wherein The specific method for evaluating the entity recognition performance of the trained model using the test dataset is as follows: Evaluate the trained model using the labeled test data, and calculate the performance metrics of the model in entity recognition. The performance metrics include accuracy, recall, and F1 score. The accuracy is the ratio of the number of entities correctly predicted by the model to the total number of entities; the recall is the ratio of the number of words correctly predicted as entities to the number of words of real entities; the F1 score is the harmonic mean that comprehensively considers precision and recall.
8. An earthquake prevention and disaster reduction entity recognition system that integrates attention and MarkBERT, characterized in that, It includes: Data processing module: used to collect the original text data in the field of earthquake in a specific period and preprocess the collected original text data; Feature extraction module: used to label the preprocessed text data according to the preset annotation objectives and rules, and extract features from the labeled data; Model training module: used to build a deep learning model including a MarkBERT encoding layer, a time series analysis layer, a dynamic context adjustment module, and a CRF decoding layer, and train the model using the labeled data; Model testing module: used to evaluate the entity recognition performance of the trained model using the test dataset.
9. An electronic device, characterized in that, It includes: Processor; And, A memory arranged to store computer-executable instructions, which when executed cause the processor to implement the steps of the earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, Used to store computer-executable instructions, which when executed implement the steps of the earthquake prevention and disaster reduction entity recognition method integrating attention and MarkBERT as described in any one of claims 1 to 7.
Citation Information
Cited By
Earthquake emergency plan named entity identification method and system
CN121168456A