Semantic conflict identification method and device based on event element extraction

Through the event feature extraction model based on multi-level label dependency modeling, combined with ERNIE 3.0 and Bagging ensemble learning, the problems of insufficient event feature extraction accuracy and weak semantic conflict recognition ability are solved, and efficient and accurate information processing is achieved, which is suitable for multi-source information fusion and public opinion monitoring.

CN120633669APending Publication Date: 2025-09-12NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510791424.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies have insufficient accuracy in event element extraction, weak semantic conflict identification capabilities, and strong dependence on training data, making it difficult to meet the high-quality information processing needs in areas such as multi-source information fusion, public opinion monitoring, and news reporting.

Method used

An event feature extraction model based on multi-level label dependency modeling is adopted, combined with the ERNIE 3.0 semantic encoding layer, label-aware attention layer, dynamic conditional random field and bagging ensemble learning, to achieve high-precision extraction of event features and semantic conflict identification through data enhancement and similarity analysis.

Benefits of technology

It significantly improves the accuracy and robustness of event element extraction, reduces dependence on training data, improves the reliability and accuracy of semantic conflict identification, and adapts to the application needs of multi-source information fusion and public opinion monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633669A_ABST
    Figure CN120633669A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic conflict recognition method and device based on event element extraction, and the method comprises the steps: carrying out the event detection of an input text, recognizing whether an event exists in the text, and determining the type of the event; determining an event element mode according to the event type; performing event element extraction on the event element mode corresponding to the text by adopting a pre-training model to obtain a labeling sequence of event elements; aligning events from different sources, and mapping the events to the same event framework after successful matching; performing semantic similarity calculation on event elements in the successfully matched multi-source event pairs; and judging whether the event elements have semantic conflicts or not according to a preset similarity threshold. According to the method, the accuracy and reliability of semantic conflict recognition can be improved, and the requirements for high-quality information processing in the fields of multi-source information fusion, public opinion monitoring, news reports and the like are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more particularly to a semantic conflict identification method and device based on event element extraction. Background Art

[0002] In the digital age, the explosive growth of text data presents unprecedented challenges for information processing and analysis. From news reports, social media, corporate documents, to academic research, massive amounts of text contain a wealth of event information. Accurately extracting and analyzing this event information is crucial for understanding content, supporting decision-making, and achieving information fusion. However, with the continuous expansion of application scenarios, particularly in multi-source information fusion and cross-domain event analysis, semantic conflict has become increasingly prominent. Semantic conflict refers to inconsistent or contradictory descriptions of the same event in texts from different sources, such as differences in time, location, participants, and other factors. Such conflicts can lead to misleading information, poor decision-making, and even uncontrolled public opinion. Therefore, identifying semantic conflict has become a key issue that needs to be urgently addressed in the field of natural language processing.

[0003] Traditional event information processing methods rely primarily on manual annotation and simple rule matching. This approach is not only inefficient but also struggles to cope with the complexity and diversity of large-scale data. In recent years, with the development of deep learning technology, pre-trained models such as BERT, RoBERTa, and ERNIE3.0 have made significant progress in natural language processing tasks. By pre-training on large-scale text corpora, these models learn rich linguistic knowledge and semantic information, providing a strong technical foundation for event information extraction and analysis. However, existing methods still have the following significant shortcomings in event feature extraction and semantic conflict identification:

[0004] (1) Insufficient accuracy in event element extraction

[0005] Existing event feature extraction methods are mostly based on traditional machine learning models or simple pre-trained models. These methods struggle to effectively capture the complex dependencies between event features when processing complex text data. For example, semantic connections exist between event triggers, participants, and time, but existing models often fail to accurately model these connections, resulting in errors in the extraction results. Furthermore, existing methods face high training difficulties and poor annotation results when processing large-scale data. Particularly when dealing with complex event information, the model's performance is limited, making it difficult to meet the needs of practical applications.

[0006] (2) Weak ability to identify semantic conflicts

[0007] Existing semantic conflict identification methods primarily rely on simple rule-based matching, which can be prone to misjudgments or omissions when aligning and matching multi-source events. For example, texts from different sources may describe the same event differently, but existing methods struggle to accurately identify whether these differences constitute a semantic conflict. Existing methods are insufficiently capable of identifying semantic conflicts in multi-source information fusion scenarios. In particular, inaccurate event element extraction reduces the reliability of semantic conflict identification, limiting the performance and application scenarios of information processing systems.

[0008] (3) Strong dependence on training data

[0009] The performance of existing event feature extraction and semantic conflict identification methods relies heavily on the quality and quantity of training data. If the training data is insufficient or low-quality, the model may not learn sufficient features to accurately complete the task. In some cases, certain categories in a dataset may have more samples than others, leading to a biased model toward the majority class and poorly labeled minority classes. Existing methods struggle to effectively improve the generalization and robustness of the models when faced with data imbalance, further impacting the accuracy of event feature extraction and semantic conflict identification.

[0010] Event element extraction is the foundation of semantic conflict identification, while semantic conflict identification is the ultimate goal. While the two complement each other, semantic conflict identification is sometimes more critical in practical applications because it directly impacts the accuracy and reliability of information. Therefore, a new method for efficiently and accurately extracting event elements and identifying semantic conflicts is urgently needed to meet the high-quality information processing needs of multi-source information fusion, public opinion monitoring, news reporting, and other fields. Summary of the Invention

[0011] In view of this, the present invention provides a semantic conflict identification method and device based on event element extraction, aiming to solve the problems of the existing technology in insufficient event element extraction accuracy, weak semantic conflict identification ability and strong dependence on training data, thereby improving the accuracy and reliability of semantic conflict identification and meeting the needs of high-quality information processing in fields such as multi-source information fusion, public opinion monitoring, and news reporting.

[0012] In order to achieve the above object, the present invention adopts the following technical solutions:

[0013] In a first aspect, an embodiment of the present invention provides a semantic conflict identification method based on event element extraction, comprising the following steps:

[0014] S1. Perform event detection on the input text to identify whether there is an event in the text and determine the event type;

[0015] S2. Determine an event element pattern according to the event type;

[0016] S3. Using a pre-trained model to extract event elements from the event element pattern corresponding to the text to obtain a labeled sequence of event elements;

[0017] S4. Align events from different sources and map them to the same event framework after successful matching.

[0018] S5. Calculate semantic similarity of event elements in successfully matched multi-source event pairs; and determine whether there is a semantic conflict between the event elements based on a preset similarity threshold.

[0019] Furthermore, the pre-trained model in step S3 is an event element extraction model based on multi-level label dependency modeling, which includes:

[0020] a) ERNIE 3.0 semantic encoding layer, which performs deep semantic encoding on the input text and generates token-level representation vectors;

[0021] b) Label-aware attention layer, which is used to establish attention weights between label embeddings and token-level representation vectors, and output label-enhanced feature representations;

[0022] c) Dynamic Conditional Random Field module, used to fuse the basic transfer matrix and the dynamic transfer matrix based on the [CLS] vector to generate sequence labeling results;

[0023] d) Ensemble learning module, which is used to introduce the Bagging ensemble learning strategy to aggregate the prediction results of multiple sub-models.

[0024] Furthermore, the event element extraction model uses a conditional mask language model to perform data enhancement on the sample data, and generates enhanced text that is semantically consistent with the sample data through a label embedding mechanism.

[0025] Furthermore, the label-aware attention layer is calculated as follows:

[0026]

[0027] Among them, h i is the token-level representation vector; e l is the label embedding vector; W is the learnable projection matrix; c i is to generate the label-aware context vector; softmax(·) is the normalization function.

[0028] Furthermore, the calculation process of the dynamic conditional random field module is as follows:

[0029] T dynamic =MLP(h [CLS] )

[0030] T final =T base +αT dynamic

[0031] Among them, T dynamic is the [CLS] vector h passed by ERNIE3.0 [CLS] , the multi-layer perceptron MLP generates the input-related dynamic transfer matrix; T base is the predefined basic transfer matrix, α is the dynamic weight coefficient; T final The transfer matrix adapted to the specific input text is obtained through weighted fusion.

[0032] Furthermore, the Bagging ensemble learning strategy in the ensemble learning module includes:

[0033] By randomly sampling the training data with replacement multiple times, multiple different training subsets are constructed;

[0034] Each training subset independently trains the event feature extraction model;

[0035] The label with the most occurrences is selected as the final prediction result through a voting mechanism.

[0036] Furthermore, the step S4 specifically includes:

[0037] Similarity analysis is used to match events from different sources. When the similarity between the calculated texts is greater than a first threshold, they are mapped to the same event framework.

[0038] Furthermore, the event element extraction model also uses a synonym replacement strategy to enhance the sample data to ensure the semantic consistency and label accuracy of the new samples.

[0039] In a second aspect, an embodiment of the present invention further provides a semantic conflict identification device based on event element extraction, using the semantic conflict identification method based on event element extraction as described in any embodiment of the first aspect, the device comprising:

[0040] The event detection module is used to detect events in the input text, identify whether there are events in the text, and determine the event type;

[0041] An element pattern determination module, configured to determine an event element pattern according to the event type;

[0042] An element extraction module is used to extract event elements from the event element pattern corresponding to the text using a pre-trained model to obtain a labeled sequence of event elements;

[0043] The multi-source event alignment and matching module is used to align events from different sources and map them to the same event framework after successful matching;

[0044] The semantic conflict identification module is used to calculate the semantic similarity of event elements in the successfully matched multi-source event pairs; and to determine whether there is a semantic conflict between the event elements based on a preset similarity threshold.

[0045] It can be seen from the above technical solutions that compared with the prior art, the present invention has the following technical advantages:

[0046] (1) Improve the precision and accuracy of event element extraction

[0047] The event factor extraction model based on multi-level label dependency modeling used in the present invention effectively captures the complex dependencies between event factors and improves the prediction accuracy of the model through the four-level collaborative architecture of ERNIE 3.0 semantic encoding layer, label perception attention layer, dynamic conditional random field (CRF) and bagging ensemble learning. The bagging algorithm is used to implement ensemble learning, further improving the accuracy and reliability of the event factor extraction task. By constructing multiple sub-models and integrating them, the generalization ability and robustness of the model can be effectively improved.

[0048] In addition, a conditional mask language model is introduced to improve the BERT model, forming Conditional Bert to implement a data enhancement method based on label embedding, improve the diversity and balance of the original data, and further improve the accuracy of event element extraction.

[0049] (2) Improve the reliability and accuracy of semantic conflict identification

[0050] This paper uses event detection technology to classify events and determine their element patterns. A proposed model extracts event elements based on these patterns. Based on the alignment and matching of multiple source events, similarity analysis is used to identify semantic conflicts. This combination of event element extraction technology and multi-source event alignment and matching effectively identifies semantic conflicts in texts from different sources, improving the reliability and accuracy of semantic conflict identification.

[0051] (3) Reduce dependence on training data

[0052] By combining a pre-trained model with a dynamic conditional random field algorithm, this paper can better learn the characteristics and dependencies of event elements with limited training data, thereby improving the model's generalization ability. Ensemble learning is achieved through the bagging algorithm, further improving the model's performance in data imbalance situations and reducing its dependence on training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0054] Figure 1 This is a flow chart of the semantic conflict identification method based on event element extraction provided by the present invention.

[0055] Figure 2 This is a schematic diagram of the semantic conflict identification method based on event element extraction provided by the present invention.

[0056] Figure 3 This is a diagram of the pre-training model architecture provided by the present invention.

[0057] Figure 4 This is a schematic diagram of the principle of integrating the learning module into the pre-training model provided by the present invention.

[0058] Figure 5 This is a block diagram of the semantic conflict identification device based on event element extraction provided by the present invention. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] The embodiment of the present invention discloses a semantic conflict identification method based on event element extraction, referring to Figure 1-2 As shown, it includes the following steps S1 to S5:

[0061] S1. Perform event detection on the input text to identify whether there is an event in the text and determine the event type;

[0062] The goal of event detection is to identify events in unstructured text. This is achieved by identifying event trigger words and classifying event types. Trigger words are words or phrases in the text that directly cause an event to occur. Event detection technology analyzes the vocabulary, grammatical structure, and semantic information in the text to identify the presence of events and classify them into different types, such as natural disasters, social events, and economic activities. For example, in news reports, event detection technology can identify event types such as "earthquakes" and "traffic accidents." Event detection is the foundation of semantic conflict identification. Accurate event classification provides precise guidance for subsequent feature extraction and conflict identification.

[0063] S2. Determine an event element pattern according to the event type;

[0064] Determine the event's element pattern based on the event type: Different event types have different element patterns. For example, natural disasters may include elements such as the disaster type, affected area, and time of the disaster, while social events may include elements such as the event subject, event behavior, and event time. Predefined event element patterns provide guidance for subsequent event element extraction. For example, for an "earthquake" event, its element pattern may include elements such as "earthquake magnitude," "earthquake occurrence time," and "earthquake affected area." These element patterns can help the model better extract key information about the event.

[0065] S3. Using a pre-trained model to extract event elements from the event element pattern corresponding to the text to obtain a labeled sequence of event elements;

[0066] The proposed pre-trained model extracts event elements based on event patterns. Text is input into the pre-trained model, which tags each word based on the event pattern and trained parameters, thereby extracting information such as the event trigger, participants, and time. Accurate event element extraction is crucial for semantic conflict identification. Only by extracting accurate event elements can effective semantic conflict identification be performed based on multi-source event alignment and matching.

[0067] In this step, the pre-trained model is an event feature extraction model based on multi-level label dependency modeling. Its framework significantly improves the accuracy and robustness of event feature extraction by deeply integrating the semantic understanding ability of the pre-trained language model with the structural constraints between labels. Figure 3As shown in the figure, the model adopts a hierarchical and progressive design concept and consists of the following four key components. The innovations of this multi-layered architecture lie in: 1) the label attention layer enables a smooth transition from the semantic space to the label space; 2) the dynamic CRF retains the advantages of sequence modeling while also gaining contextual awareness; and 3) the components form a synergistic effect, with ERNIE 3.0 responsible for semantic understanding, the label attention layer focusing on label associations, the dynamic CRF ensuring sequence legitimacy, and ensemble learning improving overall robustness.

[0068] 1. ERNIE 3.0 semantic encoding layer

[0069] ERNIE3.0 is a pre-trained deep neural network model that integrates self-supervised learning and knowledge fusion mechanisms. It adopts the Transformer architecture and integrates autoregressive and autoencoding network designs, showing strong language understanding capabilities in the field of natural language processing. Based on the semantic encoding layer of ERNIE 3.0 as the basic feature extractor, it uses its powerful Transformer architecture to perform deep semantic encoding on the input text and generate a token-level representation vector h with rich contextual information. i This stage focuses on global semantic understanding of text and capturing local features, laying the foundation for subsequent label dependency modeling.

[0070] 2. Label-aware attention layer

[0071] The Label-Aware Attention Layer is innovatively introduced, which defines a trainable high-dimensional embedding vector e for each BIO label (such as B-Person, I-Time, etc.) l , a label semantic space is established. For each token’s hidden state h i , calculate its attention weight with all label embeddings, and generate the label-aware context vector c i Specifically, the process interacts through a learnable projection matrix W and finally outputs the splicing feature [h i ;c i ].

[0072]

[0073] This design achieves three major advantages:

[0074] 1) Explicitly establish the association between tokens and potential labels, making the model label-aware at the feature level;

[0075] 2) Dynamically adjust the importance of labels through the attention mechanism, avoiding the limitation of traditional CRF in treating all labels equally and alleviating the overfitting of CRF in dependence on local labels;

[0076] 3) It provides the subsequent CRF layer with feature representation enhanced by label information, which greatly reduces the difficulty of CRF learning label transfer rules.

[0077] 3. Dynamic Conditional Random Field Module

[0078] The Dynamic Conditional Random Field (Dynamic CRF) module is an improvement to the traditional CRF. This module uses the basic transfer matrix Tbase of CRF to retain the traditional CRF's modeling ability for the legitimacy of label sequences. On this basis, it innovatively generates the input-related dynamic transfer matrix T through the [CLS] vector of ERNIE3.0. dynamic Finally, the transfer matrix T adapted to the specific input text is obtained through weighted fusion final .

[0079] T dynamic =MLP(h [CLS] )

[0080] T final =T base +αT dynamic

[0081] This design enables the model to automatically adjust the strength of label transfer constraints based on the text type (such as formal news or social media) and event category (such as natural disasters or social events) to adapt to different event types. For example, the continuity requirement of the temporal element is strengthened when processing news reports, while this restriction is appropriately relaxed in social media text.

[0082] 4. Integrated learning module

[0083] Ensemble learning based on event element extraction model: In order to further improve the accuracy and reliability of event element extraction tasks, this paper introduces the Bagging algorithm to implement ensemble learning. Figure 4As shown in the figure, the bagging algorithm is based on a self-service sampling and aggregation technique. It constructs multiple training subsets by performing multiple random samplings with replacement on the training data. Each training subset independently trains a corresponding base model, and the prediction results of these sub-models are integrated. Before each new model is trained, independent sampling with replacement is performed again to ensure the uniqueness of the training data for each base model. During the prediction phase, each sentence is tested on different base models to obtain multiple sets of prediction results. A voting mechanism is then used to select the BIO label with the highest number of votes for each word, constructing a complete annotation sequence and identifying event elements. For example, 10 sub-models are constructed through multiple random samplings. Each sub-model is trained on independent sampled data. Ultimately, a voting mechanism selects the label with the highest number of occurrences as the final prediction result, significantly improving the model's generalization and robustness.

[0084] In this method, although individual base models may exhibit bias due to using only partial data for training, multi-model voting can effectively reduce the variance of the ensemble model and significantly improve the model's stability and prediction accuracy. Because the training processes of each model are independent of each other, the ensemble model is less dependent on the training data and exhibits greater adaptability when processing unknown test data, thus avoiding the limitations of a single model and enhancing generalization capabilities. Furthermore, the integration characteristics of the Bagging algorithm effectively suppress the interference of noisy data. Even if individual models are affected by noise and misjudge, the prediction results of other models can be corrected to maintain the stability of the overall model performance. Through the random sampling of sample data by the Bagging algorithm and the integration of multiple models, both the accuracy and robustness of the event factor extraction model are improved.

[0085] In this embodiment, the model integrates the pre-trained language model and the label-aware attention mechanism, realizes context-adaptive label transfer modeling through the dynamic conditional random field algorithm, significantly improves the sequence labeling accuracy, and combines the Bagging ensemble learning strategy to enhance the model generalization ability, thereby comprehensively optimizing the accuracy and robustness of event element extraction.

[0086] In addition, in the tasks of event element extraction and semantic conflict identification, data sets often face the dual challenges of diversified event types and uneven sample distribution, which poses a significant obstacle to model training. In order to overcome this problem, this embodiment introduces the Conditional BERT technology based on the Conditional Mask Language Model (C-MLM) to implement data enhancement, so as to improve the model's detection efficiency for various events. C-MLM randomly masks some tokens in the input sentence, and combines contextual semantics with sentence label information to predict the original vocabulary index of the masked word, thereby generating a new type of text that is compatible with the label semantics. Compared with the prediction mechanism of the traditional masked language model that only relies on contextual information, C-MLM innovatively uses label embeddings to replace BERT's native sentence embeddings, deeply integrating event label information into the model prediction process. This improvement enables the model to accurately predict masked words based on the dual information of contextual context and label semantics, and then generate new text samples that are strictly consistent with the original sentence labels.

[0087] In addition, Conditional Bert further expands the sample data scale through a synonym replacement strategy. In specific implementation, the model uses a label embedding mechanism to ensure that the replaced sentences are highly consistent with the original sentences at the semantic level, while strictly maintaining the accuracy of the labels. This data augmentation process not only effectively enriches the sample diversity of the dataset, but also significantly alleviates the problem of uneven sample distribution, providing higher-quality training data support for the event feature extraction task. Through this enhancement strategy that combines semantic consistency and label accuracy, Conditional Bert significantly improves the model's generalization ability and prediction accuracy in event feature extraction tasks during the data augmentation process.

[0088] S4. Align events from different sources and map them to the same event framework after successful matching.

[0089] In a multi-source information fusion scenario, there may be multiple texts describing the same event, but these texts may differ in terms of expression, detailed information, etc. Therefore, these multi-source events need to be aligned for subsequent semantic conflict identification. Event alignment is the basis of multi-source event fusion, and accurate event alignment can provide accurate input for subsequent semantic conflict identification. Similarity analysis is used to match multi-source events. The similarity analysis method can determine whether they describe the same event by calculating the similarity between texts. The cosine similarity calculation method can be used to calculate the similarity value between texts based on the feature vectors of event elements. Cosine similarity is an indicator used to measure the degree of similarity between two vectors in direction, and its value range is between -1 and 1. The calculation formula for cosine similarity is as follows:

[0090]

[0091] Where A and B are two vectors, A·B represents the dot product of vectors A and B, and ||A|| and ||B|| represent the moduli of vectors A and B, respectively.

[0092] Specifically, the dot product of vectors A and B is calculated as:

[0093]

[0094] Among them, A i and B i are the components of vectors A and B in the i-th dimension, and n is the dimension of the vector.

[0095] The formula for calculating the modulus of vector A is:

[0096]

[0097] The closer the cosine similarity value is to 1, the more similar the directions of the two vectors are; the closer the value is to -1, the more opposite the directions of the two vectors are; and a value of 0 indicates that the two vectors are orthogonal, meaning their directions are completely unrelated. When the similarity value exceeds a certain threshold, the two events are considered a match, and further semantic conflict identification can be performed. For example, by calculating the cosine similarity of two event descriptions, if the similarity value is greater than 0.8, the two events are considered a match and subsequent semantic conflict identification can be performed.

[0098] S5. Calculate semantic similarity of event elements in successfully matched multi-source event pairs; and determine whether there is a semantic conflict between the event elements based on a preset similarity threshold.

[0099] For successfully matched multi-source event pairs, semantic conflicts are identified for event elements. A semantic conflict occurs when the descriptions of the same event element in text from different sources are inconsistent or contradictory. For example, if one text describes the time element of an event as "January 1, 2023" while another describes it as "January 2, 2023," this constitutes a semantic conflict. Accurate conflict detection is the core of semantic conflict identification. Only by accurately detecting conflicts can appropriate resolution strategies be implemented.

[0100] Similarly, similarity analysis is used to compare the semantics of event elements. For each event element, its descriptions from different source texts are extracted, and the semantic similarity between them is calculated. If the semantic similarity falls below a certain threshold, a semantic conflict is considered to exist. Word embedding technology can be used to convert the descriptions of event elements into vector representations, and then the similarity between the vectors is calculated to identify semantic conflicts. For example, by calculating the cosine similarity of the time elements of two event descriptions, if the similarity value is less than 0.5, the two time elements are considered to have a semantic conflict.

[0101] The present invention provides a semantic conflict identification method based on event element extraction, which adopts an event element extraction model of multi-level label dependency modeling to realize the extraction of event elements. The model integrates the pre-trained language model with the label-aware attention mechanism, realizes context-adaptive label transfer modeling through the dynamic conditional random field algorithm, and enhances the generalization ability of the model by combining the Bagging ensemble learning strategy, which significantly improves the accuracy and reliability of event element extraction. In addition, during the training process, sample expansion is achieved based on label embedding, and semantically consistent data enhancement is achieved through the label embedding mechanism. Finally, in terms of semantic conflict identification, event detection technology is used to complete event classification and element pattern determination, event elements are extracted through the proposed pre-training model, and on the basis of multi-source event alignment and matching, the similarity analysis method is used to realize accurate identification of semantic conflicts. The present invention has broad application prospects in the fields of multi-source information fusion, public opinion monitoring, news reporting, etc., and can significantly improve the efficiency and accuracy of information processing.

[0102] Based on the same inventive concept, the embodiment of the present invention further provides a semantic conflict identification device based on event element extraction, using the semantic conflict identification method based on event element extraction as any of the above embodiments, referring to Figure 5 As shown, the device includes:

[0103] The event detection module is used to detect events in the input text, identify whether there are events in the text, and determine the event type;

[0104] An element pattern determination module, configured to determine an event element pattern according to the event type;

[0105] An element extraction module is used to extract event elements from the event element pattern corresponding to the text using a pre-trained model to obtain a labeled sequence of event elements;

[0106] The multi-source event alignment and matching module is used to align events from different sources and map them to the same event framework after successful matching;

[0107] The semantic conflict identification module is used to calculate the semantic similarity of event elements in the successfully matched multi-source event pairs; and to determine whether there is a semantic conflict between the event elements based on a preset similarity threshold.

[0108] This device achieves full-process optimization, from data enhancement to semantic conflict identification, through multi-layered technological innovations. In terms of core model architecture, an event feature extraction model based on multi-level label dependency modeling is constructed. This model significantly improves the accuracy and robustness of feature extraction through a four-level collaborative architecture consisting of the ERNIE 3.0 semantic encoding layer, label-aware attention layer, dynamic conditional random fields, and bagging ensemble learning. Specifically, in label dependency modeling, the innovative introduction of a label-aware attention mechanism and a dynamic CRF transfer matrix enables a smooth transition from semantic space to label space. Furthermore, during model training, a Conditional BERT technique based on the Conditional Masked Language Model (C-MLM) is designed for data preprocessing. This achieves semantically consistent data enhancement through a label embedding mechanism, effectively addressing sample imbalance. At the application level, a confidence-weighted multi-source event alignment method and a hierarchical semantic conflict identification strategy are proposed. Through feature-level similarity analysis and dynamic thresholding, high-precision matching and conflict detection across source events are achieved, providing an efficient and reliable technical solution for multi-source information fusion, public opinion monitoring, and other fields.

[0109] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0110] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A semantic conflict identification method based on event element extraction, characterized in that: The steps include: S1. Perform event detection on the input text to identify whether there is an event in the text and determine the event type; S2. Determine an event element pattern according to the event type; S3. Using a pre-trained model to extract event elements from the event element pattern corresponding to the text to obtain a labeled sequence of event elements; S4. Align events from different sources and map them to the same event framework after successful matching. S5. Calculate semantic similarity of event elements in the successfully matched multi-source event pairs; Determine whether there is a semantic conflict between event elements based on the preset similarity threshold.

2. The semantic conflict identification method based on event element extraction according to claim 1 is characterized in that: The pre-trained model in step S3 is an event element extraction model based on multi-level label dependency modeling, which includes: a) ERNIE 3.0 semantic encoding layer, which performs deep semantic encoding on the input text and generates token-level representation vectors; b) Label-aware attention layer, which is used to establish attention weights between label embeddings and token-level representation vectors, and output label-enhanced feature representations; c) Dynamic Conditional Random Field module, used to fuse the basic transfer matrix and the dynamic transfer matrix based on the [CLS] vector to generate sequence labeling results; d) Ensemble learning module, which is used to introduce the Bagging ensemble learning strategy to aggregate the prediction results of multiple sub-models.

3. The semantic conflict identification method based on event element extraction according to claim 2 is characterized in that: The event element extraction model uses a conditional mask language model to perform data enhancement on sample data, and generates enhanced text that is semantically consistent with the sample data through a label embedding mechanism.

4. The semantic conflict identification method based on event element extraction according to claim 2 is characterized in that: The label-aware attention layer is calculated as follows: Among them, h i is the token-level representation vector; e l is the label embedding vector; W is the learnable projection matrix; c i is to generate the label-aware context vector; softmax(·) is the normalization function.

5. The semantic conflict identification method based on event element extraction according to claim 2 is characterized in that: The calculation process of the dynamic conditional random field module is as follows: T dynamic =MLP(h [CLS] ) T final =T base +αT dynamic Among them, T dynamic is the [CLS] vector h passed by ERNIE3.0 [CLS] , the multi-layer perceptron MLP generates the input-related dynamic transfer matrix; T base is the predefined basic transfer matrix, α is the dynamic weight coefficient; T final The transfer matrix adapted to the specific input text is obtained through weighted fusion.

6. The semantic conflict identification method based on event element extraction according to claim 2 is characterized in that: The Bagging ensemble learning strategy in the ensemble learning module includes: By randomly sampling the training data with replacement multiple times, multiple different training subsets are constructed; Each training subset independently trains the event feature extraction model; The label with the most occurrences is selected as the final prediction result through a voting mechanism.

7. The semantic conflict identification method based on event element extraction according to claim 1 is characterized in that: The step S4 specifically includes: Similarity analysis is used to match events from different sources. When the similarity between the calculated texts is greater than a first threshold, they are mapped to the same event framework.

8. The semantic conflict identification method based on event element extraction according to claim 2 is characterized in that: The event element extraction model also uses a synonym replacement strategy to enhance the sample data to ensure the semantic consistency and label accuracy of the new samples.

9. A semantic conflict identification device based on event element extraction, characterized in that: Using the semantic conflict identification method based on event element extraction according to any one of claims 1 to 8, the device comprises: The event detection module is used to detect events in the input text, identify whether there are events in the text, and determine the event type; An element pattern determination module, configured to determine an event element pattern according to the event type; An element extraction module is used to extract event elements from the event element pattern corresponding to the text using a pre-trained model to obtain a labeled sequence of event elements; The multi-source event alignment and matching module is used to align events from different sources and map them to the same event framework after successful matching; The semantic conflict identification module is used to calculate the semantic similarity of event elements in the successfully matched multi-source event pairs; and to determine whether there is a semantic conflict between the event elements based on a preset similarity threshold.

Citation Information

Cited By

  • Emergency judgment method based on knowledge graph

    CN122220521A