Multi-modal event extraction method for non-paired single-modal labeling
Through the combination of dynamic mask module and fine-grained modal alignment module, the semantic correlation inaccurate of multimodal event extraction under non-paired single-modal annotation is solved, efficient multimodal event extraction is achieved, and the model's generalization ability in real scenarios is improved.
Patent Information
- Application Number
- CN202510352953.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, under unpaired single-modal annotated data, the multimodal event extraction method has problems such as inaccurate cross-modal semantic correlation and poor generalization performance of synthetic data, resulting in insufficient generalization ability of the model in real scenarios.
The dynamic mask module is used to adaptively filter the low confidence noise proposal in the synthetic data, and a fine-grained modal alignment module is introduced to set a learnable query vector for each category label, and modal fusion is carried out through the self-attention mechanism to achieve fine-grained modeling of event structured elements.
It effectively suppresses semantic distortion in synthetic data, improves the robustness and accuracy of the model on real multimodal data, realizes low-cost and high-precision multimodal event extraction, and supports applications such as public opinion analysis and news summary.
Smart Images

Figure CN120296108A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multimodality and information extraction, and particularly relates to a multimodal event extraction task in a weakly supervised scenario of unpaired unimodal annotation, which can be applied to scenarios such as news summarization and public opinion analysis. Background Art
[0002] Event extraction aims to quickly and accurately identify events occurring in unstructured data and jointly extract structured event information, including event types, participating entities, etc., which can greatly improve the efficiency of downstream tasks such as public opinion analysis and news summarization. With the rapid development of Internet technology, the amount of multimodal data containing images and texts has increased explosively. Therefore, how to extract event information from multimodal data has become an urgent task to be solved. Traditional multimodal methods rely on paired multimodal annotation data to establish semantic correlation modeling between modalities through cross-modal alignment mechanisms such as contrastive learning. However, this method requires high-quality and paired multimodal event extraction annotation data for model training, and in real scenarios, it requires high labor and time costs for annotation work.
[0003] To reduce the dependence on data, existing technologies propose to use only unpaired unimodal annotation data during training, that is, text annotation data and image annotation data without corresponding relationships. Existing technologies use a generative model to generate synthetic data of another modality for unimodal annotation data, construct a pseudo-multimodal training set, and thus perform information alignment between modalities. However, this alignment focuses on the overall correlation between modalities but ignores the fine-grained modeling and alignment of the internal structured elements of events required for event extraction, resulting in inaccurate cross-modal semantic correlation. In addition, the synthetic data may have hallucinations or artifacts different from real data, leading to the model relying on the synthetic data distribution and having poor generalization performance on real multimodal data.
[0004] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] To solve the above technical problems, the present invention proposes a multimodal event extraction method for unpaired unimodal annotation data.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect of the present invention, a multimodal event extraction method for unpaired unimodal annotation includes:
[0008] S1. Based on the unpaired single-modal annotation data of news waiting to extract events, use text-to-image and image-to-text models to generate synthetic data of another modality, and construct a pseudo multi-modal training set;
[0009] S2. Process the text data and image data respectively, extract the candidate words in the text and the candidate target boxes in the image as candidate proposals, and obtain the feature representations and confidences of all candidate proposals;
[0010] S3. In the training stage, use the dynamic masking module to adaptively mask the candidate proposals of the synthetic modality based on the confidence to suppress low-quality synthetic data;
[0011] S4. Set learnable query vectors for each class label, interact them with the feature representations of the candidate proposals in the sample, and extract the feature representations of the event elements corresponding to the class in the sample;
[0012] S5. Determine the class labels of the candidate proposals according to the feature similarity between the event elements and the candidate proposals of the real modality, and realize multi-modal event extraction.
[0013] Further, in step S2, when processing the text data, extracting the candidate proposals and obtaining the feature vectors and confidences specifically includes:
[0014] Represent each sentence of the text data as a token sequence W = {w1, w2,..., w n}, use the text sequence representation model to learn the sequence W, and output the class of each word. The optional classes are {O, B1, I1, B2, I2,... B m , I m}, where O means that the word does not belong to any word describing event elements, B m means that the word is the beginning of a word describing an event element belonging to class m, and I m means that the word is the subsequent part of a word describing an event element belonging to class m;
[0015] Merge each group of consecutive beginnings B m and subsequent parts I m of the same type to obtain all text candidate proposals. The feature vector of each candidate proposal is the average pooling of the output vector of the last hidden layer of the text sequence representation at the positions of the words it contains, and its confidence is the mean of the probabilities that the words it contains belong to this type.
[0016] Further, in step S2, when processing the image data, extracting the candidate proposals and obtaining the feature vectors and confidences specifically includes:
[0017] Use the object detection model to process the image to obtain all target boxes and their confidences;
[0018] Process the image using an image feature extraction model, divide the complete image into squares of a given resolution, and obtain the feature vector of each square;
[0019] Average pool the feature vectors of the squares covered by each target box as the candidate proposal feature, and the confidence level is the confidence level of the target box; when the complete image is used as a candidate proposal, its feature vector is the average pooling of all square features, and the confidence level is 1.
[0020] Furthermore, in step S3, during the training phase, use the dynamic mask module to adaptively mask the candidate proposals of the synthetic modality, specifically including:
[0021] Map the feature vectors of the image and text candidate proposals to the same dimension through linear transformation;
[0022] For each candidate proposal of the synthetic modality, calculate the cosine similarity between its feature vector and the feature vector of each candidate proposal of the corresponding real modality, and use it as the weight to calculate the weighted average of the confidence levels, and multiply it by the similarity of the synthetic modality candidate proposal to obtain the adjusted confidence level;
[0023] According to the masking ratio of the linear transformation, select the part of the candidate proposals with the lowest confidence levels in the synthetic modality for masking to exclude the negative impact of the low-confidence synthetic data features on the model; use the self-attention method to perform modality fusion of the feature vectors of the real modality and the synthetic modality candidate proposals, and obtain the candidate proposals H of all modalities = Encoder([X T , X I ), where Encoder represents the encoder structure in Transformer, and X T and X I respectively represent the feature vectors of the text and image candidate proposals.
[0024] Furthermore, in step S4, the feature extraction of the event elements specifically includes: setting a learnable query vector for each class label, and interacting with the feature representation of the candidate proposals in the sample to extract the feature representation of the event elements corresponding to the class in the sample, where:
[0025] Preset a series of classes Y = {Y0, Y1, Y2, …, Y n}, and each class Y i corresponds to a learnable query vector q i , and the query vector learns the feature representation corresponding to the class during the training process;
[0026] Input the class query vector q and the candidate proposal H into the decoder Decoder of Transformer, and calculate the feature vector Z of the event elements corresponding to the class = Decoder(q, H, H).
[0027] Further, in step S5, according to the feature similarity between the event elements and all real modal candidate proposals, the class labels of the candidate proposals are determined to achieve multi-modal event extraction, specifically including:
[0028] Calculate the feature vector Z of each event element i And the cosine similarity with the feature vector H of each candidate proposal j
[0029] Use the softmax function to calculate the similarity weights of H j For all categories of event elements of Z i These weights respectively represent the probabilities that the candidate proposal belongs to different event element categories p = Softmax(cos(H, Z));
[0030] During training, use this probability to calculate the cross-entropy loss of the model;
[0031] In the inference stage, use the similarity threshold to determine whether the candidate proposal belongs to the corresponding event element category.
[0032] In the second aspect of the present invention, a multi-modal event extraction system for unpaired single-modal annotation includes:
[0033] A text-to-image and image-to-text model for obtaining synthetic data of another modality according to unpaired single-modal annotation data to construct a pseudo multi-modal training set;
[0034] A text sequence representation model for processing text data, extracting candidate proposals in the text and obtaining feature representations and confidence levels;
[0035] A target detection model and an image feature extraction model for processing image data, extracting candidate proposals in the image and obtaining feature representations and confidence levels;
[0036] A dynamic masking module for adaptively masking candidate proposals of the synthetic modality during the training stage;
[0037] A fine-grained modality alignment module for setting a learnable query vector for each class label and interacting with the feature representations of candidate proposals in the sample to extract the feature representations of event elements of the corresponding class in the sample;
[0038] A candidate proposal classifier for determining the class labels of candidate proposals according to the feature similarity between event elements and all real modal candidate proposals to achieve multi-modal event extraction.
[0039] In the third aspect of the present invention, an electronic device includes:
[0040] A processing unit configured to perform data processing operations;
[0041] A storage medium communicatively connected to the processing unit and storing computer-executable instructions;
[0042] Wherein, when the computer-executable instructions are executed by the processing unit, the multi-modal event extraction method for non-paired single-modal annotation is implemented.
[0043] In a fourth aspect of the present invention, a computer-readable storage medium stores a computer program, and the method is implemented when the computer program is executed by a processor.
[0044] In a fifth aspect of the present invention, a computer program product includes a computer program, and the method is implemented when the computer program is executed by a processor.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] The present invention proposes a multi-modal event extraction method for non-paired single-modal annotation. By using a dynamic mask module to adaptively filter out low-confidence noise proposals in synthetic data, it effectively suppresses semantic distortion and artifact interference generated by the generation model and reduces the model's dependence on the synthetic data distribution. At the same time, a fine-grained modality alignment module is introduced to set learnable query vectors for each event category label, enabling it to interact with candidate proposal features and explicitly model the cross-modal associations of event structured elements (such as event types, participating entities), solving the problem of inaccurate semantic associations caused by global alignment in traditional methods. A pseudo multi-modal training set is generated using the unpaired text dataset ACE2005 and image dataset imSitu, and verified on the real multi-modal dataset M2E2. As Figure 3 shown, the proposed method achieves significant advantages in the multimedia event extraction task, demonstrates higher robustness compared to existing methods, realizes low-cost and high-precision joint extraction of structured event information in the weakly supervised scenario of single-modal annotation, and provides efficient technical support for applications such as public opinion analysis and news summary. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the multi-modal event extraction method for non-paired single-modal annotation of the present invention.
[0048] Figure 2 is a schematic diagram of the processing flow of real image annotation data in the multi-modal event extraction method for non-paired single-modal annotation according to an embodiment of the present invention;
[0049] Figure 3 is the quantitative evaluation result of the multi-modal event extraction method for non-paired single-modal annotation according to an embodiment of the present invention.
[0050] Figure 4 Schematic diagram of a multi-modal event extraction system for unpaired single-modal annotation according to an embodiment of the present invention. Detailed implementation manners
[0051] In order to make the technical problems, technical solutions and beneficial effects to be solved by the embodiments of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Refer to Figure 1 , an embodiment of the present invention provides a multi-modal event extraction method for unpaired single-modal annotation, including:
[0053] S1. Based on the unpaired single-modal annotation data of events waiting to be extracted from news, use the text-to-image and image-to-text models to generate synthetic data of another modality, and construct a pseudo multi-modal training set;
[0054] S2. Process the text data and image data respectively, extract the candidate words in the text and the candidate target boxes in the image as candidate proposals, and obtain the feature representations and confidences of all candidate proposals;
[0055] S3. In the training stage, use the dynamic masking module to adaptively mask the candidate proposals of the synthetic modality based on the confidence to suppress low-quality synthetic data;
[0056] S4. Set a learnable query vector for each class label, make it interact with the feature representations of the candidate proposals in the sample, and extract the feature representations of the event elements corresponding to the class in the sample;
[0057] S5. Determine the class labels of the candidate proposals according to the feature similarity between the event elements and the candidate proposals of the real modality, and realize multi-modal event extraction.
[0058] The method for multi-modal event extraction for unpaired single-modal annotation according to the embodiments of the present invention first uses a generative model to obtain synthetic data of another modality based on the unpaired single-modal annotation data of the event to be extracted, and constructs a pseudo multi-modal training set; processes text and image data respectively, extracts candidate proposals in the text to obtain feature representations and confidence levels; in the training stage, uses a dynamic masking module to adaptively mask the candidate proposals of the synthetic modality, so as to filter out low-confidence synthetic data with hallucinations and artifacts; uses a fine-grained modality alignment module to set a learnable query vector for each class label and interact with the feature representations of the candidate proposals, extracts the feature representations of the corresponding classes, and realizes fine-grained modality alignment based on the structured elements of the event; determines the class labels of the candidate proposals according to the feature similarities between the event elements and all real modality candidate proposals, and realizes multi-modal event extraction. The present invention realizes low-cost and high-robust multi-modal event extraction in the weakly supervised scenario of single-modal annotation, and can provide technical support for event-related tasks such as news summarization and public opinion analysis.
[0059] The following further describes the method for multi-modal event extraction for unpaired single-modal annotation proposed in the preferred embodiments of the present invention in conjunction with specific examples.
[0060] As Figure 1 shown, the method for multi-modal event extraction for unpaired single-modal annotation according to the embodiments of the present invention includes:
[0061] S1. Based on the unpaired single-modal annotation data, use text-to-image and image-to-text models to obtain synthetic data of another modality, and construct a pseudo multi-modal training set;
[0062] S2. Process text and image data respectively, extract candidate words in the text and candidate target boxes in the image (collectively referred to as candidate proposals), and obtain the feature representations and confidence levels of all candidate proposals;
[0063] S3. In the training stage, use a dynamic masking module to adaptively mask the candidate proposals of the synthetic modality;
[0064] S4. Set a learnable query vector for each class label, and interact with the feature representations of the candidate proposals in the sample, and extract the feature representations of the event elements of the corresponding classes in the sample;
[0065] S5. Determine the class labels of the candidate proposals according to the feature similarities between the event elements and all real modality candidate proposals, and realize multi-modal event extraction.
[0066] Figure 2 The processing flow of the real image annotation data in this embodiment is given.
[0067] In this embodiment, for unpaired single-modal annotation data, the steps of using text-to-image and image-to-text models to obtain synthetic data of another modality and constructing a pseudo-multi-modal training set include:
[0068] S11. Obtain synthetic image data of real text annotation data;
[0069] S12. Obtain synthetic text data of real image annotation data.
[0070] In this embodiment, the real text annotation data consists of sentences in the ACE2005 dataset and their corresponding event types and event argument labels. The diffusion model Stable Diffusion v2.1 is used to generate 5 images with a resolution of 512×512 for each sentence in the ACE20005 dataset, and 100 steps of denoising are used.
[0071] In this embodiment, the real text annotation data consists of images in the imSitu dataset and their corresponding scenarios and entity labels. The BLIP model is used to generate 1 text description for each original image in imSitu, and the nucleus sampling method with a probability threshold of 0.9 is used during generation.
[0072] In this embodiment, the steps of processing text and image data, extracting candidate proposals, and obtaining feature vectors and confidences include:
[0073] S21. Obtain the feature representation sequence of the target text;
[0074] S22. Obtain the candidate proposals of the target text and their feature vectors and confidences;
[0075] S23. Obtain the target box of the target image;
[0076] S24. Obtain the candidate proposals of the target image and their feature vectors and confidences.
[0077] In this embodiment, each sentence of the text data is represented as a token sequence W = {w1, w2, …, w n}, and there are m categories of event elements. BERT is used to learn the sequence W and output the feature representation of each word:
[0078]
[0079] where X = x1, x2, …, x n is the hidden state vector of the last layer of BERT, and d is the dimension of the hidden layer.
[0080] Next, for m types of event elements, the words can be divided into 2m + 1 categories: {O, B1, I1, B2, I2, … B m , Im}, where O indicates that the word does not belong to any word describing event elements, and B i indicates that the word is the beginning of a word describing an event element belonging to the i-th category, and I i indicates that the word is the subsequent part of a word describing an event element belonging to the i-th category. Use a linear classification head and the Softmax function to calculate the probabilities belonging to each category based on the hidden state vector of the word:
[0081]
[0082] where P T = p1, p2, …, p n . Take the one with the maximum probability as the category of the word, and combine each group of consecutive beginnings and subsequent parts of the same type to obtain all text candidate proposals.
[0083]
[0084] The feature vector of each candidate proposal is the average pooling of the output vector of the last hidden layer of the text sequence representation at the positions of the words it contains, and its confidence is the mean of the probabilities that the words it contains belong to this type. For example, for the candidate proposal T k = [i s , i e , its feature vector X k and confidence Conf(T k ) are respectively:
[0085]
[0086] In this embodiment, use the object detection model YOLOv8 to process the image to obtain all object bounding boxes and their confidences. Use a 12-layer CLIP model to process the image, divide the complete image into 16×16 grids, and obtain the feature vector of each grid. The combination of all grids covered by each object bounding box is used as a candidate proposal for the image, and its feature vector is the average pooling of the feature vectors of all grids covered by this object bounding box, and the confidence is the confidence of the object bounding box. In addition, the complete image is also used as a candidate proposal, and its feature vector is the average pooling of the feature vectors of all grids, and the confidence is 1.
[0087] X k = AvgPool(CLIP(patches(Ik)))
[0088] In this embodiment, in the training stage, the steps of using the dynamic mask module to adaptively mask the candidate proposals of the synthetic modality include:
[0089] S31. Map the feature vectors of the image and text candidate proposals to the same dimension using linear transformation;
[0090] S32. Adjust the confidence of the synthetic modality candidate proposals according to the similarity to the ground truth modality candidate proposals;
[0091] S33. Dynamically mask the synthetic modality candidate proposals according to the confidence and training progress;
[0092] S34. Fuse the ground truth modality candidate proposals and the unmasked synthetic modality candidate proposals.
[0093] In this embodiment, the candidate proposals of the image and text are mapped to vectors of 768 dimensions. For each candidate proposal of the synthetic modality, calculate the cosine similarity of its feature vector with each candidate proposal of the corresponding ground truth modality, and use it as the weight to calculate the weighted average of the confidence, and multiply it by the similarity of the synthetic modality candidate proposal to obtain the adjusted confidence:
[0094]
[0095] where R represents the proposal of the ground truth modality, S represents the proposal of the synthetic modality, X represents the proposal feature, Conf(·) represents the initial confidence of the proposal, represents the adjusted confidence of the synthetic modality.
[0096] In this embodiment, in the training stage, set the masking rate to decrease from 50% to 20%, and select the part of the candidate proposals with the lowest confidence in the synthetic modality for masking, aiming to exclude the negative impact of the low-confidence synthetic data features on the model:
[0097]
[0098] where t total is the total number of training epochs, t is the current training epoch, α init and α final represent the start and end masking rates respectively, α(t) is the current masking rate, is the current masked synthetic modality proposal, and Quantile represents the quantile.
[0099] After that, use the self-attention method to perform modality fusion on the feature vectors of the ground truth modality and the masked synthetic modality candidate proposals to obtain the candidate proposals of all modalities:
[0100]
[0101] where X T and X IThey respectively refer to the features of text and image proposals, where N and M are the numbers of text and image proposals. Emcpder represents the encoder structure in the Transformer.
[0102] In this embodiment, the step of setting a learnable query vector for each class label and interacting with the feature representations of candidate proposals in the sample to extract the feature representations of event elements corresponding to the class in the sample includes:
[0103] S41. Set the class query vector;
[0104] S42. Obtain the feature vectors of event elements corresponding to the class according to the class query vector and the candidate proposals.
[0105] In this embodiment, a series of classes Y = {Y0, Y1, Y2, …, Y m} are preset in advance, and each class Y i corresponds to a learnable query vector q i . These query vectors will learn the feature representations corresponding to the classes during the training process. Input the class query vector q and the candidate proposals H into the decoder Decoder of the Transformer to calculate the feature vectors of event elements corresponding to the class:
[0106] Z = Decoder(q, H, H)
[0107] In this embodiment, the step of determining the class labels of candidate proposals according to the feature similarity between event elements and all true-modal candidate proposals to achieve multi-modal event extraction includes:
[0108] S51. Calculate the feature similarity between event elements and candidate proposals;
[0109] S52. Calculate the probabilities that the candidate proposals belong to different event element classes.
[0110] In this embodiment, calculate the cosine similarity between the feature vector of each event element and the feature vector of each candidate proposal. Use the softmax function to calculate the similarity ratios of the proposals to event elements of all classes, and these ratios respectively represent the probabilities that the candidate proposals belong to different event element classes:
[0111] P(H) = Softmax(cos(H, Z))
[0112] In this embodiment, in the training stage, the cross-entropy loss of the probability calculation model of the true-modal proposal r:
[0113]
[0114] In the inference stage, a similarity threshold is used to determine whether the candidate proposal belongs to the corresponding event element category.
[0115] Example verification
[0116] The multi-modal event extraction method for unpaired single-modal annotation according to the embodiment of the present invention uses the text event extraction dataset ACE2005 with single-modal annotation and the image event extraction dataset imSitu for training, and is tested and verified on the multi-modal event extraction dataset M2E2. Figure 3 For the quantitative evaluation results of the multi-modal event extraction method for unpaired single-modal annotation according to the embodiment of the present invention. As Figure 3 shown, this method has achieved significant advantages in the multimedia event extraction task (the ED / EAE indicators reach 58.7 / 34.3), showing higher robustness than the comparison models CLIP-Event (53.4 / 23.4), MGIM (55.6 / 24.6), etc., and realizing the joint extraction of low-cost and high-precision structured event information in a weakly supervised scenario, providing efficient technical support for applications such as public opinion analysis and news summary.
[0117] As Figure 4 shown, the present invention also provides a multi-modal event extraction system for unpaired single-modal annotation, which implements the multi-modal event extraction method for unpaired single-modal annotation as described above, including:
[0118] A text-to-image and image-to-text model for obtaining synthetic data of another modality based on unpaired single-modal annotation data to construct a pseudo multi-modal training set;
[0119] A text sequence representation model for processing text data, extracting candidate proposals in the text and obtaining feature representations and confidences;
[0120] A target detection model and an image feature extraction model for processing image data, extracting candidate proposals in the image and obtaining feature representations and confidences;
[0121] A dynamic masking module for adaptively masking candidate proposals of the synthetic modality in the training stage;
[0122] A fine-grained modality alignment module for setting a learnable query vector for each category label and interacting with the feature representations of candidate proposals in the sample to extract the feature representations of event elements of the corresponding category in the sample;
[0123] A candidate proposal classifier for determining the category label of the candidate proposal according to the feature similarity between the event element and all real modality candidate proposals, and realizing multi-modal event extraction.
[0124] The embodiment of the present invention also provides an electronic device, including:
[0125] A processing unit configured to perform data processing operations;
[0126] A storage medium communicatively connected to the processing unit and storing computer-executable instructions;
[0127] Wherein, when the computer-executable instructions are executed by the processing unit, the processing unit implements the steps of the multimodal event extraction method for unpaired unimodal annotation as described above.
[0128] An embodiment of the present invention also provides a storage medium for storing a computer program, which when executed performs at least the method as described above.
[0129] An embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is configured to perform at least the method as described above when executing the computer program.
[0130] An embodiment of the present invention also provides a processor that executes a computer program and performs at least the method as described above.
[0131] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, Ferromagnetic Random Access Memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0132] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0133] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0134] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0135] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.
[0136] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention essentially or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks or optical disks and other various media that can store program codes.
[0137] In the method disclosed in several method embodiments provided by the present invention, they can be arbitrarily combined without conflict to obtain new method embodiments.
[0138] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0139] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0140] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious modifications can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.
Claims
1. A multi-modal event extraction method for unpaired single-modal annotation, characterized in that, Including: S1. Based on the unpaired single-modal annotation data of the event to be extracted, use the text-to-image and image-to-text models to generate synthetic data of another modality, and construct a pseudo multi-modal training set; S2. Process the text data and image data respectively, extract the candidate words in the text and the candidate target boxes in the image as candidate proposals, and obtain the feature representations and confidences of all candidate proposals; S3. In the training stage, use the dynamic masking module to adaptively mask the candidate proposals of the synthetic modality based on the confidence to suppress low-quality synthetic data; S4. Set a learnable query vector for each class label, make it interact with the feature representation of the candidate proposals in the sample, and extract the feature representation of the event elements corresponding to the class in the sample; S5. Determine the class label of the candidate proposal according to the feature similarity between the event element and the candidate proposal of the real modality, and realize multi-modal event extraction.
2. The multimodal event extraction method for unpaired single-modal annotation according to claim 1, wherein In step S2, when processing the text data, extracting the candidate proposals and obtaining the feature vectors and confidences specifically includes: Each sentence of the text data is represented as a token sequence W = {w1, w2, …, w n}, and the text sequence representation model is used to learn the sequence W and output the category of each word. The optional categories are {O, B1, I1, B2, I2, … B m , I m}, where O indicates that the word does not belong to any word describing event elements, and B m indicates that the word is the beginning of a word describing an event element belonging to category m, and I m indicates that the word is the subsequent part of a word describing an event element belonging to category m; Combine each group of consecutive starting Bs belonging to the same type m with the subsequent part I m to obtain all text candidate proposals. The feature vector of each candidate proposal is the average pooling of the output vectors of the last hidden layer of the text sequence representation at the word positions it contains, and its confidence is the mean of the probabilities that the words it contains belong to this type.
3. The multimodal event extraction method for unpaired single-modal annotation according to claim 1, wherein In step S2, when processing the image data, extracting the candidate proposals and obtaining the feature vectors and confidences specifically includes: Use the object detection model to process the image to obtain all the target boxes and their confidences; Use the image feature extraction model to process the image, divide the complete image into squares with a given resolution, and obtain the feature vector of each square; The feature vectors of the squares covered by each target box are average-pooled as the candidate proposal features, and the confidence is the confidence of the target box; when the complete image is used as the candidate proposal, its feature vector is the average-pooling of all the square features, and the confidence is 1.
4. The multi-modal event extraction method for unpaired single-modal annotation according to claim 1, wherein In step S3, in the training stage, using the dynamic masking module to adaptively mask the candidate proposals of the synthetic modality specifically includes: Map the feature vectors of the image and text candidate proposals to the same dimension through linear transformation; For each candidate proposal of the synthetic modality, calculate the cosine similarity between its feature vector and the feature vector of each candidate proposal of the corresponding real modality, and use it as the weight to calculate the weighted average of the confidences, and multiply it by the similarity with the candidate proposal of this synthetic modality to obtain the adjusted confidence; According to the linearly varying mask ratio, select the part of the candidate proposals with the lowest confidence in the synthetic modality for masking to exclude the negative impact of low-confidence synthetic data features on the model; use self-attention to perform modality fusion of the candidate proposal feature vectors of the real modality and the synthetic modality to obtain the candidate proposals H of all modalities = Encoder([X T , X I ), where Encoder represents the encoder structure in Transformer, and X T and X I represent the candidate proposal feature vectors of text and image respectively.
5. The multimodal event extraction method for unpaired unimodal annotation according to claim 1, characterized in that In step S4, the feature extraction of the event elements specifically includes: setting a learnable query vector for each class label, and interacting with the feature representation of the candidate proposals in the sample to extract the feature representation of the event elements corresponding to the class in the sample, where: A series of categories Y = {Y0, Y1, Y2, …, Y n} are preset, and each category Y i corresponds to a learnable query vector q i , and the query vector learns the feature representation of the corresponding category during the training process; Input the class query vector q and the candidate proposal H into the decoder Decoder of the Transformer, and calculate the feature vector Z of the event elements corresponding to the class = Decoder(q, H, H).
6. The multimodal event extraction method for unpaired unimodal annotation according to claim 1, wherein, In step S5, according to the feature similarity between the event element and all the candidate proposals of the real modality, determine the class label of the candidate proposal to realize multi-modal event extraction, specifically including: Calculate the feature vector Z of each event element i with each candidate proposed feature vector H j for cosine similarity Calculate H using the softmax function j For Z i The similarity weights of the event elements for all categories, where these weights represent the probabilities p that the candidate proposal belongs to different event element categories respectively: p = Softmax(cos(H, Z)); During training, use the cross-entropy loss of this probability calculation model; In the inference stage, use the similarity threshold to determine whether the candidate proposal belongs to the corresponding event element class.
7. A multi-modal event extraction system for non-paired single-modal annotation, characterized in that, Including: The text-to-image and image-to-text models are used to obtain synthetic data of another modality according to the unpaired single-modal annotation data and construct a pseudo multi-modal training set; A text sequence representation model for processing text data, extracting candidate proposals in the text, and obtaining feature representations and confidences; An object detection model and an image feature extraction model for processing image data, extracting candidate proposals in the image, and obtaining feature representations and confidences; A dynamic masking module for adaptively masking candidate proposals of the synthetic modality during the training phase; A fine-grained modality alignment module for setting a learnable query vector for each class label and interacting with the feature representations of candidate proposals in the sample to extract the feature representations of event elements corresponding to the class in the sample; A candidate proposal classifier for determining the class label of the candidate proposal according to the feature similarity between the event element and all true modality candidate proposals, thereby implementing multi-modal event extraction.
8. An electronic device, characterized in that, The electronic device includes: A processing unit configured to perform data processing operations; A storage medium communicatively connected to the processing unit and storing computer-executable instructions; Wherein, when the computer-executable instructions are executed by the processing unit, the multi-modal event extraction method for unpaired single-modal annotation as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1 to 6.