Information extraction method, device, electronic device and storage medium for multi-view learning
Through the multi-perspective learning information extraction method, the multi-headed attention mechanism is used for interactive capture and classification, which solves the problems of inconsistency in the information source and incompatibility of targets faced by existing models in the emotional event extraction task, and significantly improves the accuracy and robustness of the task.
Patent Information
- Application Number
- CN202510289568.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing models face inconsistent information source, insufficient semantic understanding, and incompatibility with sentiment analysis and event extraction goals in the emotional event extraction task, resulting in failure to achieve the expected results.
The information extraction method of multi-view learning is adopted, and the initial text is obtained, and the encoder is set to perform semantic capture, which is converted into context features and semantic representation. Then, the multi-head attention mechanism is used to capture context features and type features interactively, output interactive features, and input classifiers for event classification and emotion classification respectively to generate extraction results.
It significantly improves the accuracy and robustness of the emotional event extraction task, enhances the accuracy of emotional polarity judgment, and better captures event-related information, improving overall performance.
Smart Images

Figure CN119782533B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an information extraction method, device, electronic device and storage medium for multi-view learning. Background Art
[0002] With the rapid development of information technology, social media, news websites, etc. have become important platforms for obtaining and disseminating information. This information not only contains a large amount of event information, but also contains emotional tendencies related to the event. Event extraction and sentiment analysis tasks play a vital role in practical applications, such as public opinion monitoring, market analysis, crisis management, etc.
[0003] However, current extraction models face a series of challenges in the task of emotional event extraction. These problems make the performance of existing models in the task of emotional event extraction fail to achieve the expected results.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those skilled in the art. Summary of the invention
[0005] In view of this, the present application proposes an information extraction method, device, electronic device and storage medium for multi-view learning to solve or partially solve the above problems.
[0006] Based on the above objectives, this application provides an information extraction method for multi-view learning, including:
[0007] Acquire an initial text, and perform semantic capture on the initial text by setting an encoder to obtain context features and semantic representation;
[0008] The semantic representation is converted according to a preset classification type to obtain a type feature; wherein the preset classification type includes at least event classification and sentiment classification;
[0009] Interactively capturing the context feature and the type feature according to a multi-head attention mechanism, and outputting an interactive feature;
[0010] Inputting the interaction features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification;
[0011] According to the first prediction result and the second prediction result, a corresponding information template is determined, and the information template is filled and decoded according to the context feature to generate an extraction result.
[0012] In some exemplary embodiments, after obtaining the initial text, the method further includes:
[0013] The initial text is preprocessed for data cleaning according to preset rules.
[0014] In some exemplary embodiments, before setting the encoder to perform semantic capture on the initial text, the method further includes:
[0015] Use the text segmentation algorithm to extract keywords from the preprocessed initial text, perform word embedding conversion on the extracted keywords, and obtain word embedding vectors;
[0016] The word embedding vector is input into the set encoder.
[0017] In some exemplary embodiments, after outputting the interaction feature, the method further includes:
[0018] Mean pooling is performed on the interaction features to enhance the capture capability of the interaction features.
[0019] In some exemplary embodiments, filling and decoding the information template according to the context feature includes:
[0020] According to the prompt label of the information template, information is extracted from the context feature and spliced with the information template;
[0021] The spliced information template is input into the set decoder for data decoding.
[0022] In some exemplary embodiments, the prompt label is generated by a natural language template and processed by a linear transformation and an activation function.
[0023] In some exemplary embodiments, the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification are trained by joint optimization, and the loss function used in the training is composed of the weighted sum of the loss of event classification and the loss of sentiment classification.
[0024] Based on the same concept, the present application also provides an information extraction device for multi-view learning, comprising:
[0025] The first module is used to obtain an initial text, and to capture the semantics of the initial text by setting an encoder to obtain context features and semantic representation;
[0026] The second module is used to convert the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type at least includes event classification and sentiment classification;
[0027] The third module is used to interactively capture the context feature and the type feature according to a multi-head attention mechanism, and output an interactive feature;
[0028] A fourth module is used to input the interaction features into a classifier corresponding to the event classification and a classifier corresponding to the sentiment classification, respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification;
[0029] The fifth module is used to determine the corresponding information template according to the first prediction result and the second prediction result, fill and decode the information template according to the context feature, and generate an extraction result.
[0030] Based on the same concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the methods described above when executing the program.
[0031] Based on the same concept, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement any of the methods described above.
[0032] Based on the same concept, the present application also provides a computer program product, including computer program instructions, which, when executed on a computer, enable the computer to execute any of the above methods.
[0033] As can be seen from the above, the present application provides a method, device, electronic device and storage medium for information extraction of multi-perspective learning, the method comprising: obtaining an initial text, capturing the semantics of the initial text by setting an encoder, obtaining context features and semantic representation; converting the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type includes at least event classification and sentiment classification; interactively capturing context features and type features according to a multi-head attention mechanism, and outputting interactive features; inputting the interactive features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification, respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification; according to the first prediction result and the second prediction result, determining the corresponding information template, filling and decoding the information template according to the context features, and generating an extraction result. The present application can effectively extract event information and its related sentiment information at the same time through a multi-head attention mechanism, and establish a strong correlation between the two, thereby significantly improving the accuracy and robustness of the sentiment event extraction task. This method can not only enhance the accuracy of sentiment polarity judgment, but also better capture event-related information, thereby improving overall performance. Furthermore, the multi-head attention mechanism can effectively capture the interactive features between the contextual features in the initial text and the event type embedding. This mechanism significantly improves the performance of event type judgment and sentiment polarity classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 A flowchart of an exemplary method provided in an embodiment of the present application.
[0036] Figure 2 A schematic diagram of the structure of an exemplary device provided in an embodiment of the present application.
[0037] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of this specification more clear, this specification is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0039] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the examples of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "including" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0040] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0041] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0042] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0043] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0044] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0045] In order to make the purpose, technical solutions and advantages of this specification more clear, this specification is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0046] As mentioned in the background technology section, the event extraction task aims to extract event information from the context, including event type, trigger words and arguments; the sentiment analysis task needs to identify the sentiment polarity in the text. In the related technology, event extraction methods can be mainly divided into pipeline models and joint reasoning models. The pipeline model usually extracts trigger words first and then extracts arguments. Although it is simple to implement and intuitive to operate, it is prone to error accumulation, that is, the error of the previous task will affect the accuracy of subsequent tasks. In order to solve the problem of error accumulation, the joint reasoning model came into being. It can extract trigger words and arguments at the same time, thereby reducing the impact of error propagation, but for complex event extraction tasks, there are still certain challenges, especially when deep context understanding is required.
[0047] As for sentiment analysis tasks, in related technologies, sentiment analysis methods can be divided into explicit sentiment analysis and implicit sentiment analysis. Explicit sentiment analysis can usually be judged directly based on the sentiment expression in the text, while implicit sentiment analysis faces more challenges because the sentiment information is usually not expressed explicitly but hidden in the context. Although existing sentiment analysis and event extraction methods have achieved certain results in their respective fields, the sentiment event extraction task that combines the two still faces two major challenges: inconsistent information sources and incompatible research objectives.
[0048] In practical applications, there are significant differences between the information published on news websites and social media. The information on news websites is usually factual, while social media is full of users' emotional expressions. Since factual information and emotional information are expressed in different forms and frequencies in text, this brings about the problem of inconsistent information sources for the task of emotional event extraction. To address this challenge, some existing methods attempt to integrate events and their emotional polarity in news and comments by constructing multimodal datasets, but this is still a problem that needs further optimization.
[0049] In addition, the event extraction task requires not only the identification of event types, but also the extraction of corresponding trigger words and arguments, while the sentiment analysis task focuses on the overall sentiment polarity of the text. Due to the differences in their goals, the research objectives are incompatible. Therefore, how to effectively combine sentiment analysis and event extraction tasks and achieve joint optimization of the two has become a major challenge in the current research on sentiment event extraction.
[0050] It can be seen from the above that the event argument extraction model of the method in the above embodiment faces a series of challenges in the task of sentiment event extraction, mainly including the inconsistency of information sources, the lack of semantic understanding, and the incompatibility between sentiment analysis and event extraction goals. These problems make the performance of the existing model in the task of sentiment event extraction fail to achieve the expected effect.
[0051] In view of the above practical situation, the embodiment of the present application provides an information extraction method for multi-perspective learning. The present application can effectively extract event information and its related emotional information at the same time through a multi-head attention mechanism, and establish a strong correlation between the two, thereby significantly improving the accuracy and robustness of the emotional event extraction task. This method can not only enhance the accuracy of emotional polarity judgment, but also better capture event-related information, thereby improving the overall performance. Furthermore, the use of a multi-head attention mechanism can effectively capture the interactive features between the contextual features in the initial text and the event type embedding. This mechanism significantly improves the performance of event type judgment and emotional polarity classification tasks.
[0052] Figure 1 A flow chart of an exemplary method provided in an embodiment of the present application is shown.
[0053] like Figure 1 As shown, the information extraction method of multi-view learning proposed by an embodiment of the present application exemplarily includes the following steps.
[0054] Step 102, obtaining an initial text, and setting an encoder to perform semantic capture on the initial text to obtain context features and semantic representation.
[0055] In this step, the initial text is the initial data used for information extraction. The text here can be obtained by the user through the corresponding port when using the corresponding smart device, or it can be imported in batches through tables or links. After that, the initial text needs to be encoded to generate a context feature representation. The initial text can generally be composed of several words, and then each word in the text can be firstly converted into a word embedding, and the word embedding vector is passed into the corresponding encoder as input. The encoder processes the text based on these embedding vectors, captures the grammatical and semantic information in the text, and finally outputs context features (H is used to represent the context features in this embodiment). After that, the initial text is converted into a word embedding and converted into a vector representation. Then, these word embeddings are further encoded through the encoder to obtain the semantic representation E of each initial text.
[0056] In some more specific scenarios, the initial text can be obtained from news comments and public data sets. Afterwards, in order to further improve the data quality of the initial text, the data can be preprocessed according to the event type, etc., and irrelevant data can be removed to ensure that the data used meets the standardization requirements of the event type. Then, the event data is cleaned, for example, irrelevant information such as HTML tags can be deleted to obtain standardized Chinese event data. That is, in some embodiments, after obtaining the initial text, the method also includes: performing data cleaning preprocessing on the initial text according to preset rules.
[0057] Afterwards, when training the model formed by the method of the present application, event data containing sentiment annotations can be collected from news websites and social media websites. Afterwards, the data quality can be further improved by preprocessing the collected data, which may include removing text content irrelevant to the event, deleting HTML tags, and screening based on recommendations from domain experts.
[0058] In some embodiments, in order to facilitate the acquisition of contextual features, a text segmentation algorithm can also be used to extract keywords of the event, and filter out content related to the event. In a more specific scenario, the text segmentation algorithm can be a TextRank algorithm. Afterwards, for each segmented word, a word embedding conversion is performed, and the word embedding vector is passed into the encoder as input. The encoder processes the text based on these embedding vectors, captures the grammatical and semantic information in the text, and finally outputs contextual features. In a more specific scenario, the set encoder can be a BART encoder encoder, that is, when performing sentiment analysis on each event type, the initial text is first encoded using the BART pre-trained model to obtain the contextual feature representation H of the text. That is, in some embodiments, before the initial text is semantically captured by setting the encoder, the method also includes: extracting keywords from the pre-processed initial text using a text segmentation algorithm, performing word embedding conversion on the extracted keywords, and obtaining a word embedding vector; and inputting the word embedding vector into the set encoder.
[0059] In more specific scenarios, different event types can be pre-selected for the initial text, such as political, economic, and technical categories, etc. Afterwards, each major category can be further classified into several sub-event types based on different emotional expressions. For example, the emotional polarity can be defined as:
[0060]
[0061] After that, the BART model can be used to encode the context and represent the input text ,in Different keywords represent the initial text. After passing through the BART encoder, we get the context embedding: ,in, is the dimension of the hidden state, express After converting the event type text into word embeddings, we can continue to obtain the semantic representation of each event type by encoding the hidden state: ,in, is the number of event types, is the maximum length of the event type text, is the dimension of the hidden state, express dimensional matrix.
[0062] Step 104, converting the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type at least includes event classification and sentiment classification.
[0063] In this step, after obtaining the semantic representation Afterwards, in order to obtain the final event type embedding representation , that is, the final type feature is obtained . It can be expressed in semantics On the basis of the above, it passes through a fully connected layer and converts it into a type feature that is more conducive to model calculation. . For each event type Perform the following calculations: ,in, and are learnable parameters.
[0064] Afterwards, the preset classification types can be set according to the explanation in the aforementioned step 102, that is, different event types can be pre-selected and set first, for example, they can be divided into political, economic, and technical categories, etc., and then, for each major category, several sub-event types can be further classified according to different emotional expressions, for example, emotional polarity can be defined as containing positive emotions, containing negative emotions, no clear emotional tendency, etc. Among them, the classification of event types corresponds to event classification, and the subsequent classification of emotional types corresponds to emotional classification.
[0065] Finally, the type characteristics obtained in this step are and the context features obtained in step 102 It is the basis for the model to understand the emotional information and event structure in the text, and can be used as the input of subsequent models.
[0066] Step 106: Interactively capture the context feature and the type feature according to the multi-head attention mechanism, and output the interactive feature.
[0067] In this step, the multi-head attention mechanism is mainly used to model the interaction between context features and type features. With type traits The relationship between them is deeply modeled to capture the complex interaction information between the two, thereby obtaining richer relationship information between context and event type. This process generates interaction features , which fully reflects the dynamic interaction between context and event type. In specific application scenarios, a multi-head attention mechanism is used to With type traits The processing formula is expressed as .
[0068] In this step, the multi-head attention mechanism uses multiple parallel attention calculations to make the context features With type traits Able to interact in different subspaces to generate richer type features in event type embedding representation .
[0069] In some embodiments, in order to further enhance the model's ability to capture global context information, the interactive features Based on this, mean pooling is further performed to generate a global interactive feature representation Through the pooling operation, the key information in the interactive features can be effectively integrated to obtain a more compact and representative global feature representation. Specifically, the updated interactive feature representation This representation brings together the global information of sentiment analysis and event type features, providing more comprehensive semantic support for the prediction of sentiment polarity and event type. That is, in some embodiments, after outputting the interaction features, the method further includes: performing mean pooling on the interaction features to enhance the capture capability of the interaction features.
[0070] Step 108, inputting the interaction features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification.
[0071] In this step, the interaction features Or updated global features The input is sent to two independent classifiers to predict the event type and sentiment polarity respectively. In this process, the prediction results of event type (corresponding to the first prediction result) and the prediction results of sentiment polarity (Corresponding to the second prediction result) can be calculated through the SoftMax function to obtain the respective probability distribution, thereby obtaining the final prediction output.
[0072] In a more specific scenario, the softmax classifier is used to utilize global features Predict the event type and sentiment polarity. Specifically, the prediction result of event type And the prediction results of sentiment polarity Calculated by the following formula:
[0073]
[0074] in, and is a learnable parameter, and is the bias term. Finally, and They represent the predicted probabilities of event type and sentiment polarity respectively. This process ensures accurate prediction of sentiment polarity and event type by processing global context information by the classifier.
[0075] Step 110, determining a corresponding information template according to the first prediction result and the second prediction result, filling and decoding the information template according to the context feature, and generating an extraction result.
[0076] In this step, after determining the first prediction result And the second prediction result Afterwards, the corresponding information template can be determined based on the two results, and then the prompt-guided decoding method can be used to generate the extraction result. In a specific scenario, a corresponding information template can be designed and created for each event type, that is, for different first prediction results And the second prediction result Each event type is designed with a corresponding information template. Each information template contains information about a specific event type and the argument roles associated with it. These templates will be input into the model to obtain the extraction results of event trigger words and various argument roles. In this process, the model will combine the input context information and the designed information templates for processing to generate accurate extraction results. Finally, the model will output the specific start and end positions of each event argument based on the guidance of the context information and information templates. This process helps to improve the accuracy and precision of event argument extraction.
[0077] In more specific scenarios, for prompts designed for fine-grained event extraction, for example, for a "conflict" event, a prompt is constructed as "[Conflict]: [time], at [location], [attacker] clashed with [victim], involving a total of [number] people." By constructing an information template for each event type, combining the event type information and trigger words, a complete information template is formed and input into the model to assist in extracting the event arguments.
[0078] Afterwards, the contextual features of the original text can be embedded and combined with the determined information template and input into the BART decoder for further encoding and decoding. Through this process, the model can learn the prompt information of event trigger words and argument roles and generate accurate event argument extraction results.
[0079] In some embodiments, in order to embed and splice information more accurately, after designing a special natural language information template for each event type, the context features and prompts can be embedded in the template. and tag embedding The information template is filled and decoded according to the context features, including: extracting information from the context features according to the prompt tag of the information template, and splicing it with the information template; inputting the spliced information template into the set decoder for data decoding. The prompt tag can be the aforementioned prompt embedding and tag embedding , then the decoder can be comparable to the previous encoder, for example both are BART codecs, etc.
[0080] In a more specific application scenario, the prompt label may be generated through a natural language template, the template includes a description of the event type, the argument role and its corresponding entity type, and the prompt is processed through a linear transformation and an activation function, thereby generating complete event information in the decoder. That is, in some embodiments, the prompt label is generated through a natural language template and processed through a linear transformation and an activation function.
[0081] Finally, the generated extraction results can be output, for example, they can be displayed on a corresponding device to give the operator corresponding feedback. Of course, in other embodiments, the output method of the extraction results may not be limited to output display, but can also be used to store, display, use or reprocess the extraction results. According to different application scenarios and implementation needs, the specific output method of the extraction results can be flexibly selected.
[0082] Specifically, for example, in an application scenario where the method of this embodiment is executed on a single device, the extraction results can be directly output in a displayed manner on a display component (display, projector, etc.) of the current device, so that the operator of the current device can directly see the content of the extraction results from the display component.
[0083] For another example, for an application scenario where the method of this embodiment is executed on a system composed of multiple devices, the extraction results can be sent to other preset devices as receivers in the system, that is, synchronization terminals, through any data communication method (wired connection, NFC, Bluetooth, wifi, cellular mobile network, etc.), so that the synchronization terminal can perform subsequent processing on them. Optionally, the synchronization terminal can be a preset server, which is generally set up in the cloud as a data processing and storage center, which can store and distribute the extraction results; wherein the recipient of the distribution is the terminal device, and the holder or operator of these terminal devices can be the manager, supervisor, maintenance and sales personnel of the information analysis system, etc.
[0084] For another example, in an application scenario where the method of this embodiment is executed on a system composed of multiple devices, the extraction results can be sent directly to a preset terminal device through any data communication method, and the terminal device can be one or more of the ones listed in the preceding paragraphs.
[0085] In more specific application scenarios, a corresponding model can be formed according to this method. The training process and use process of the model are similar to the above steps. Only during the training process, the training set data is annotated with prompts. During use, only the prompt part needs to be removed, and only the core context information and the prediction results generated by the model are retained to obtain the final event extraction results. By removing the prompt part, redundant information of the model when generating the final result is avoided, thereby improving the accuracy and efficiency of event extraction.
[0086] For the specific training process, training can be performed by jointly optimizing the loss functions of the sentiment polarity classification task and the event type determination task. The purpose is to enable the model to take into account the correlation between sentiment analysis and event extraction tasks, as well as their possible goal conflicts when processing sentiment analysis and event extraction tasks. Through this optimization method, the model can achieve good comprehensive performance on multiple tasks, ensuring that when performing sentiment polarity classification and event type determination, not only the accuracy of a single task is improved, but also the synergy between multiple tasks is guaranteed. The loss function is composed of the weighted sum of the sentiment classification loss and the event classification loss, and optimization is achieved by weighted combination of the losses of different tasks. The specific loss function can be expressed as ,in is the sentiment classification loss function, is the event type classification loss function, and is a weight parameter. That is to say, for the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification, joint optimization can be performed by weighted sum to ensure the joint optimization of the two tasks. That is, in some embodiments, the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification are trained by joint optimization, and the loss function used in the training is composed of the weighted sum of the loss of the event classification and the loss of the sentiment classification.
[0087] Afterwards, for the event type classification loss function More specifically, the generated event type embedding and event argument role embedding can be input into the model after feature fusion. The self-attention mechanism is used to accurately select the start and end positions of event arguments, and the model parameters are updated by optimizing the cross entropy loss function, thereby improving the accuracy of event argument extraction. The specific event type classification loss function Can be .
[0088] It can be seen from the above embodiments that the embodiment of the present application provides a method for information extraction of multi-perspective learning, the method comprising: obtaining an initial text, capturing the semantics of the initial text by setting an encoder, obtaining context features and semantic representation; converting the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type includes at least event classification and sentiment classification; interactively capturing context features and type features according to a multi-head attention mechanism, and outputting interactive features; inputting the interactive features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification, respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification; according to the first prediction result and the second prediction result, determining the corresponding information template, filling and decoding the information template according to the context features, and generating an extraction result. The present application can effectively extract event information and its related sentiment information at the same time through a multi-head attention mechanism, and establish a strong correlation between the two, thereby significantly improving the accuracy and robustness of the sentiment event extraction task. This method can not only enhance the accuracy of sentiment polarity judgment, but also better capture event-related information, thereby improving overall performance. Furthermore, the multi-head attention mechanism can effectively capture the interactive features between the contextual features in the initial text and the event type embedding. This mechanism significantly improves the performance of event type judgment and sentiment polarity classification tasks.
[0089] It can be seen that in a more specific application scenario, the embodiment of the present application first encodes the input text through an encoder and a multi-head attention mechanism, models event features and sentiment features, and determines the event type and comment sentiment polarity. Then, event information is generated using the contextual features encoded by the encoder and the prompt-guided decoding, and event trigger words and arguments are extracted through argument role prompts described in natural language. Finally, multi-task learning is achieved through a multi-head attention mechanism, and the rich semantic representations therein are used to achieve efficient extraction of sentiment and event information. In this way, the incompatibility problem between sentiment analysis and event extraction goals is effectively solved through multi-task learning, thereby significantly improving the accuracy of sentiment polarity analysis and event argument extraction.
[0090] By applying the method and corresponding model of the embodiment of the present application, event information and its related sentiment polarity can be effectively extracted at the same time, solving the problem of incompatibility between sentiment analysis and event extraction task objectives. By jointly optimizing the loss functions of these two tasks, a stronger connection can be established between sentiment polarity analysis and event extraction, thereby significantly improving the accuracy and robustness of the sentiment event extraction task. This method can not only enhance the accuracy of sentiment polarity judgment, but also better capture event-related information, thereby improving overall performance.
[0091] At the same time, multi-task learning achieved through the multi-head attention mechanism, using its framework and discriminant method, can effectively solve the incompatibility problem between sentiment analysis and event extraction tasks in practical applications. Unlike traditional generative models, the embodiments of the present application can effectively avoid the occurrence of the "hallucination" phenomenon, which usually causes the generative model to make wrong predictions or generate irrelevant information when processing complex tasks. Furthermore, the use of a multi-head attention mechanism can effectively capture the interactive information between contextual features in the text and event type embeddings. This mechanism significantly improves the performance of event type judgment and sentiment polarity classification tasks. By combining the manually designed prompt information with the contextual features extracted by the BART encoder, the model can accurately extract event trigger words and arguments, ensuring the efficient and precise execution of sentiment event extraction tasks.
[0092] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of the embodiment of the present application can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.
[0093] It should be noted that the above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an information extraction device for multi-perspective learning.
[0095] refer to Figure 2 , the information extraction device for multi-view learning comprises:
[0096] The first module 210 is used to obtain an initial text, and to perform semantic capture on the initial text by setting an encoder to obtain context features and semantic representation.
[0097] The second module 220 is used to convert the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type at least includes event classification and sentiment classification.
[0098] The third module 230 is used to interactively capture the context feature and the type feature according to the multi-head attention mechanism, and output the interactive feature.
[0099] The fourth module 240 is used to input the interaction features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification, respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification.
[0100] The fifth module 250 is used to determine the corresponding information template according to the first prediction result and the second prediction result, fill and decode the information template according to the context feature, and generate an extraction result.
[0101] In some exemplary embodiments, the first module 210 is further configured to:
[0102] The initial text is preprocessed for data cleaning according to preset rules.
[0103] In some exemplary embodiments, the first module 210 is further configured to:
[0104] Use the text segmentation algorithm to extract keywords from the preprocessed initial text, perform word embedding conversion on the extracted keywords, and obtain word embedding vectors;
[0105] The word embedding vector is input into the set encoder.
[0106] In some exemplary embodiments, the third module 230 is further configured to:
[0107] Mean pooling is performed on the interaction features to enhance the capture capability of the interaction features.
[0108] In some exemplary embodiments, the fifth module 250 is further configured to:
[0109] According to the prompt label of the information template, information is extracted from the context feature and spliced with the information template;
[0110] The spliced information template is input into the set decoder for data decoding.
[0111] In some exemplary embodiments, the prompt label is generated by a natural language template and processed by a linear transformation and an activation function.
[0112] In some exemplary embodiments, the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification are trained by joint optimization, and the loss function used in the training is composed of the weighted sum of the loss of event classification and the loss of sentiment classification.
[0113] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0114] The device of the above embodiment is used to implement the corresponding multi-view learning information extraction method in the above embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0115] Based on the same concept, corresponding to any of the above-mentioned embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the information extraction method of multi-perspective learning as described in any of the above embodiments is implemented.
[0116] Figure 3 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0117] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0118] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0119] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0120] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0121] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0122] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0123] The electronic device of the above embodiment is used to implement the corresponding multi-view learning information extraction method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0124] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the information extraction method of multi-perspective learning as described in any of the above embodiments.
[0125] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0126] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the information extraction method of multi-view learning as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0127] Based on the same concept, corresponding to any of the above-mentioned embodiments, the present application also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the information extraction method for multi-view learning. Corresponding to the execution subject corresponding to each step in each embodiment of the information extraction method for multi-view learning, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0128] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the information extraction method of multi-view learning as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0129] A person skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0130] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power / ground connections to the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device may be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0131] Although the present application has been described in conjunction with specific embodiments of the present application, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the discussed embodiments.
[0132] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A multi-view learning information extraction method, characterized in that: include: Acquire an initial text, and perform semantic capture on the initial text by setting an encoder to obtain context features and semantic representation; The semantic representation is converted according to a preset classification type to obtain a type feature; wherein the preset classification type includes at least event classification and sentiment classification; Interactively capturing the context feature and the type feature according to a multi-head attention mechanism, and outputting an interactive feature; Inputting the interaction features into the classifier corresponding to the event classification and the classifier corresponding to the sentiment classification respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification; According to the first prediction result and the second prediction result, a corresponding information template is determined, and the information template is filled and decoded according to the context feature to generate an extraction result.
2. The method according to claim 1, characterized in that After obtaining the initial text, the method further includes: The initial text is preprocessed for data cleaning according to preset rules.
3. The method according to claim 1, characterized in that Before setting the encoder to semantically capture the initial text, the method further includes: Use the text segmentation algorithm to extract keywords from the preprocessed initial text, perform word embedding conversion on the extracted keywords, and obtain word embedding vectors; The word embedding vector is input into the set encoder.
4. The method according to claim 1, characterized in that: After outputting the interaction features, the method further includes: Mean pooling is performed on the interaction features to enhance the capture capability of the interaction features.
5. The method according to claim 1, characterized in that The filling and decoding of the information template according to the context feature includes: According to the prompt label of the information template, information is extracted from the context feature and spliced with the information template; The spliced information template is input into the set decoder for data decoding.
6. The method according to claim 5, characterized in that The prompt label is generated through a natural language template and processed through a linear transformation and an activation function.
7. The method according to claim 1, characterized in that The classifier corresponding to the event classification and the classifier corresponding to the sentiment classification are trained by joint optimization, and the loss function used in the training is composed of the weighted sum of the loss of event classification and the loss of sentiment classification.
8. An information extraction device for multi-view learning, characterized in that: include: The first module is used to obtain an initial text, and to capture the semantics of the initial text by setting an encoder to obtain context features and semantic representation; The second module is used to convert the semantic representation according to a preset classification type to obtain a type feature; wherein the preset classification type at least includes event classification and sentiment classification; The third module is used to interactively capture the context feature and the type feature according to a multi-head attention mechanism, and output an interactive feature; A fourth module is used to input the interaction features into a classifier corresponding to the event classification and a classifier corresponding to the sentiment classification, respectively, to obtain a first prediction result corresponding to the event classification and a second prediction result corresponding to the sentiment classification; The fifth module is used to determine the corresponding information template according to the first prediction result and the second prediction result, fill and decode the information template according to the context feature, and generate an extraction result.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-label sentiment classification method using collaborative neural network chain
CN113222059A
Short text object sentiment classification method based on multilevel interactive attention mechanism
CN113268592A