Multi-modal Event Extraction Method, Device, Equipment, Storage Medium and Program Product

The method addresses the challenge of event extraction from complex multimodal data by employing text and image processing techniques to align and match textual and visual entities, enhancing the accuracy and efficiency of event information retrieval.

CN119623616BActive Publication Date: 2025-07-15NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163955.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-15
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately extract event information from complex multimodal data due to difficulties in bridging the gap between different data modalities and understanding both global and local semantic information.

Method used

A method involving text feature extraction, image target detection, entity matching, and fine-grained alignment of textual and visual information to enhance event extraction from multimodal data, using pre-trained models like BERT-LSTM-CRF for text and YOLOv5 for image processing, followed by event abstraction with CLIP.

Benefits of technology

Enhances the accuracy and efficiency of event extraction by leveraging fine-grained multimodal alignment, effectively capturing both global and local semantic information, thereby improving the extraction of valuable event information from complex multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623616B_ABST
    Figure CN119623616B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal event extraction method, device, equipment, storage medium and program product. The method includes: extracting features from text data in multi-modal data, generating a text label sequence based on the text feature extraction result, performing object detection on image data in multi-modal data, obtaining the fine-grained text-image information of the multi-modal data based on the text label sequence and the object detection result, generating a text entity set and a visual entity set based on the fine-grained text-image information, performing entity matching between the text entity set and the visual entity set, performing fine-grained text-image alignment on the multi-modal data based on the entity matching result, and inputting the multi-modal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction, so as to realize effective analysis of global and local semantic information, thereby accurately extracting events from multi-modal data with complex semantic scenarios and effectively avoiding missing potential event values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a multimodal event extraction method, device, equipment, storage medium and program product. Background Art

[0002] In the era of big data, massive multimodal data (including text, images, audio, video and other modalities) are widely available. In particular, the rise of graphic and text social applications has increased the proportion of multimodal data, providing multimodal information extraction technology with massive data and rich application scenarios. How to efficiently obtain valuable information from multimodal data has attracted widespread attention from researchers and has become a research hotspot of great significance. Events come from cognitive science and are widely used in philosophy, linguistics, computer science and other fields. The Automatic Content Extraction Evaluation Conference defines events as: things or changes in state that occur at a specific time point or time period, within a specific geographical area, and consist of one or more actions participated by one or more roles. The task of event extraction automatically extracts event information of interest to users from unstructured or semi-structured data and expresses it in a structured form, which has far-reaching significance for people's cognition of the world.

[0003] Traditional event extraction research mainly focuses on text data, while multimodal event extraction aims to extract richer event information from multimodal data (such as graphic data) by introducing other modalities such as images as a supplement to text modality. The core of the multimodal event extraction task lies in how to overcome the differences between data of different modalities and establish cross-modal associations. Since data of different modalities usually have relatively independent semantic representations, it is difficult to establish direct connections. Although it is currently possible to extract global semantic information from text and images, the understanding of local semantic information is relatively weak. Since event information is usually more complex and requires comprehensive global and local semantics to understand, it is currently impossible to effectively and accurately extract event information from multimodal data of complex semantic scenes, and therefore it is impossible to deeply explore potential event value information. Summary of the invention

[0004] The main purpose of the present invention is to provide a multimodal event extraction method, device, equipment, storage medium and program product, aiming to solve the problem that the existing technology cannot effectively and accurately extract event information technology from multimodal data in complex semantic scenes.

[0005] To achieve the above object, the present invention provides a multimodal event extraction method, which comprises the following steps:

[0006] Extract features from the text data in the multimodal data, and generate a text label sequence based on the text feature extraction results. The text label sequence includes multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event.

[0007] Perform object detection on the image data in the multimodal data to obtain the object detection results. The object detection results include the object bounding boxes in the image data and the judgment results on whether there are visual entities in each object bounding box.

[0008] Obtain the fine-grained text-image information of the multimodal data based on the text label sequence and the object detection results, and perform entity extraction on the multimodal data based on the fine-grained text-image information to generate a text entity set and a visual entity set.

[0009] Match the text entity set with the visual entity set to obtain the entity matching results. The entity matching results include the matching results between the text entities and the visual entities belonging to the same argument role.

[0010] Perform fine-grained text-image alignment on the multimodal data based on the entity matching results, and input the multimodal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction.

[0011] Optionally, the text feature extraction results include text sequence feature information; the extracting features from the text data in the multimodal data and generating a text label sequence based on the text feature extraction results includes:

[0012] Input the text data in the multimodal data into a pre-trained language model for feature extraction to obtain the initial sequence feature information of the text data:

[0013]

[0014] where, is the initial sequence feature information, is the text sequence length, is the dimension of the word vector feature, is the text data in the multimodal data;

[0015] Input the initial sequence feature information into a pre-trained sequence feature extraction model to obtain text sequence feature information:

[0016] ,

[0017] Among them, is the text sequence feature information, is the text sequence length, is the size of the hidden layer dimension in the pre-trained sequence feature extraction model;

[0018] Generate a text label sequence based on the text sequence feature information.

[0019] Optionally, the generating a text label sequence based on the text sequence feature information includes:

[0020] Perform label annotation based on the text sequence feature information to obtain multiple text labels;

[0021] Analyze the dependency relationship between each text label, and generate a predicted label sequence based on the dependency relationship;

[0022] Calculate the sequence score of each predicted label sequence through a scoring function, where the sequence score includes an emission score and a transition score, the emission score represents the label score of each element in the predicted label sequence, and the transition score represents the score from one label to another label;

[0023] Perform normalization processing on the predicted label sequence based on the sequence score to obtain sequence probability information:

[0024]

[0025] Among them, represents the predicted label sequence, represents the text sequence length, represents the set of all possible label sequences, represents a possible label sequence, represents the scoring function, is the text sequence feature information, is the text data in the multimodal data, represents the sequence probability information, and the sequence probability information includes the predicted label sequence of the credible probability;

[0026] Determine the text label sequence from the predicted label sequence according to the sequence probability information.

[0027] Optionally, the performing object detection on the image data in the multimodal data to obtain an object detection result includes:

[0028] Perform object localization on the image data in the multimodal data to obtain the entity position information in the image data;

[0029] Annotate a target bounding box in the image data based on the entity position information;

[0030] Input the annotated image data into a pre-trained object detection model for object detection to obtain an object detection result. The training loss function of the pre-trained object detection model includes:

[0031]

[0032] where, represents the training loss function of the pre-trained object detection model, represents the bounding box loss function, and the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box. represents the confidence loss function, and the confidence loss function is used to measure the confidence that an object is contained in the target bounding box. is a binary classification loss function, and the binary classification loss function is used to determine whether a visual entity exists in the target bounding box.

[0033] Optionally, the entity matching of the text entity set and the visual entity set to obtain an entity matching result includes:

[0034] Input the text entity set and the visual entity set into a pre-trained feature encoding model for feature encoding to obtain the text feature representation information corresponding to the text entity set and the visual feature representation information corresponding to the visual entity set:

[0035]

[0036] where, is the text feature representation information of the text entity set, is the visual feature representation information of the visual entity set, is the text encoder in the pre-trained feature encoding model, is the visual encoder in the pre-trained feature encoding model, is the text entity set, is the visual entity set, is the size of the hidden layer dimension in the pre-trained feature encoding model;

[0037] Calculate the cosine similarity between each entity in the text entity set and the visual entity set based on the text feature representation information and the visual feature representation information:

[0038]

[0039] where, is the text feature representation information, is the visual feature representation information, For the text entity and the visual entity the cosine similarity between them represents the dot product of feature vectors and represents the norm of the feature vector;

[0040] Entity matching is performed according to the cosine similarity to obtain an entity matching result.

[0041] Optionally, after performing fine-grained text-image alignment on the multimodal data based on the entity matching result and inputting the fine-grained text-image aligned multimodal data into a pre-constructed event extraction model for event extraction, it further includes:

[0042] Obtain the event extraction result output by the event extraction model, where the event extraction result includes the predicted event type and predicted argument information of the multimodal data, and the predicted argument information includes argument roles, text entities, and visual entities;

[0043] Based on the predicted event type and predicted argument information, perform performance evaluation on the event extraction model to obtain a performance evaluation result, and the performance evaluation result includes:

[0044]

[0045]

[0046]

[0047] Among them, TP represents the number of entities correctly predicted, FP represents the number of entities wrongly predicted, and FN represents the number of actual entities not recognized by the event extraction model represents the model precision rate represents the model recall rate represents the harmonic mean of the model precision rate and the model recall rate;

[0048] Optimize the event extraction model according to the performance evaluation result.

[0049] In addition, to achieve the above object, the present invention also proposes a multimodal event extraction device, and the multimodal event extraction device includes:

[0050] A text event analysis module, configured to extract features from the text data in the multimodal data and generate a text label sequence based on the text feature extraction result, where the text label sequence includes multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event;

[0051] An image event analysis module for performing object detection on the image data in the multimodal data to obtain an object detection result, where the object detection result includes the object bounding boxes in the image data and the judgment results on whether there are visual entities in each object bounding box;

[0052] A fine-grained analysis module for obtaining the text-image fine-grained information of the multimodal data based on the text label sequence and the object detection result, and performing entity extraction on the multimodal data based on the text-image fine-grained information to generate a text entity set and a visual entity set;

[0053] An entity matching module for performing entity matching between the text entity set and the visual entity set to obtain an entity matching result, where the entity matching result includes the matching results between the text entities and the visual entities belonging to the same argument role;

[0054] An event extraction module for performing fine-grained text-image alignment on the multimodal data based on the entity matching result, and inputting the multimodal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction.

[0055] In addition, to achieve the above object, the present application also proposes a multimodal event extraction device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the multimodal event extraction method as described above.

[0056] In addition, to achieve the above object, the present application also proposes a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the multimodal event extraction method as described above.

[0057] In addition, to achieve the above object, the present application also provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the multimodal event extraction method as described above.

[0058] The present invention extracts features from text data in multimodal data, generates a text label sequence based on the result of text feature extraction. The text label sequence contains multiple text labels, and the text labels include the text event types of text events in the text data and the text entities corresponding to each text event. Object detection is performed on the image data in the multimodal data to obtain an object detection result. The object detection result includes the object bounding boxes in the image data and the judgment result of whether there is a visual entity in each object bounding box. The fine-grained text-image information of the multimodal data is obtained based on the text label sequence and the object detection result, and entity extraction is performed on the multimodal data based on the fine-grained text-image information to generate a text entity set and a visual entity set. The text entity set and the visual entity set are subjected to entity matching to obtain an entity matching result. The entity matching result includes the matching result between the text entity and the visual entity belonging to the same argument role. Fine-grained text-image alignment is performed on the multimodal data based on the entity matching result, and the multimodal data after fine-grained text-image alignment is input into a pre-constructed event extraction model for event extraction. Since the present invention obtains the fine-grained text-image information of multimodal data, extracts a text entity set and a visual entity set respectively based on the fine-grained text-image information, and realizes fine-grained text-image alignment based on entity matching, effectively analyzes the global and local semantic information, effectively improves the efficiency of multimodal event extraction, can accurately extract events from multimodal data with complex semantic scenarios, effectively avoids missing potential event values, and realizes accurately extracting event information in multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0060] Figure 1 It is a schematic structural diagram of a multimodal event extraction device for the hardware operating environment related to the embodiment solution of the present invention;

[0061] Figure 2 It is a schematic flowchart of the first embodiment of the multimodal event extraction method of the present invention;

[0062] Figure 3 It is a schematic flowchart of extracting text sequence feature information in an embodiment of the multimodal event extraction method of the present invention;

[0063] Figure 4 It is a schematic diagram of object detection in an embodiment of the multimodal event extraction method of the present invention;

[0064] Figure 5 It is a schematic flowchart of entity matching in an embodiment of the multi-modal event extraction method of the present invention;

[0065] Figure 6 It is a schematic flowchart of model optimization in an embodiment of the multi-modal event extraction method of the present invention;

[0066] Figure 7 It is a structural block diagram of the first embodiment of the multi-modal event extraction device of the present invention.

[0067] The realization, functional features and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0068] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] Refer to Figure 1 , Figure 1 It is a schematic structural diagram of a multi-modal event extraction device for the hardware operating environment involved in the embodiment solution of the present invention.

[0070] As Figure 1 shown, the multi-modal event extraction device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and optionally the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless-fidelity (WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0071] Those skilled in the art can understand that Figure 1 the structure shown in

[0072] As Figure 1As shown in the figure, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a multi-modal event extraction program.

[0073] In Figure 1 In the multi-modal event extraction device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the multi-modal event extraction device of the present invention may be arranged in the multi-modal event extraction device. The multi-modal event extraction device calls the multi-modal event extraction program stored in the memory 1005 through the processor 1001 and executes the multi-modal event extraction method provided by the embodiments of the present invention.

[0074] The embodiments of the present invention provide a multi-modal event extraction method. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the multi-modal event extraction method of the present invention.

[0075] In this embodiment, the multi-modal event extraction method includes the following steps:

[0076] Step S10: Extract features from the text data in the multi-modal data and generate a text label sequence based on the text feature extraction result.

[0077] It should be noted that the core of the multi-modal event extraction task lies in how to overcome the differences between different modal data and establish cross-modal associations. Since data in different modalities usually have relatively independent semantic representations, it is difficult to directly establish connections. With the progress of multi-modal pre-training technology, multi-modal pre-training models represented by the CLIP (Contrastive Language-Image Pre-training) model have demonstrated powerful text-image alignment capabilities, enabling the modeling of text and image data in a unified semantic space and realizing cross-modal information processing. However, the CLIP model still has deficiencies in understanding fine-grained concepts in complex scenarios. Specifically, the CLIP model is good at extracting the global semantic information of text and images, but is relatively weak in understanding local semantic information. Since event information is usually complex and requires comprehensive understanding of global and local semantics, it is difficult to directly apply the CLIP model to complex semantic understanding scenarios such as multi-modal event extraction.

[0078] To this end, the present embodiment performs multi-modal event extraction based on fine-grained text-image alignment, including two stages: single-modal information extraction and multi-modal information fusion. In single-modal information extraction, fine-grained event information is first extracted from images and texts using text event extraction and visual entity extraction respectively, and then fine-grained text-image alignment is performed based on the CLIP model to achieve multi-modal information fusion, and finally multi-modal event information is obtained.

[0079] It should be understood that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a computer, or a terminal electronic device that can implement the above functions. Hereinafter, taking a multi-modal event extraction device (hereinafter referred to as an extraction device) as an example, this embodiment and the following embodiments will be described.

[0080] It should be noted that the text label sequence contains multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event.

[0081] It should be noted that the multi-modal event extraction task aims to extract structured multi-modal event information from unstructured text-image data. Specifically, for a given text T and its corresponding image I, this task needs to identify all events e contained therein and their corresponding arguments a. Each event e consists of an event type and multiple arguments a. The event type is used to represent the type of the event, such as a configuration event, a detection event, etc., which are usually predefined in the dataset. A text may contain multiple events at the same time. The argument a is used to represent the entity participating in the occurrence of the event, and usually consists of an argument role , a text entity t, and a visual entity in three parts. If there is no corresponding visual entity in the image, the argument consists of only two parts: the argument role and the text entity.

[0082] In some embodiments, the extraction device can improve the BIO annotation strategy, expand the original single-label into a dual-label, so that each label can represent both the event type and the argument role at the same time, thereby extracting both events and arguments in the text in the sequence annotation task.

[0083] Further, referring to Figure 3 , Figure 3 is a schematic flow diagram of extracting text sequence feature information in an embodiment of the present invention. In order to accurately extract information from text data, the text feature extraction result includes text sequence feature information. The above step S10 may include:

[0084] Step S101: Input the text data in the multimodal data into a pre-trained language model for feature extraction to obtain the initial sequence feature information of the text data;

[0085] Step S102: Input the initial sequence feature information into a pre-trained sequence feature extraction model to obtain text sequence feature information;

[0086] Step S103: Generate a text label sequence based on the text sequence feature information.

[0087] In some embodiments, the extraction device can use a BERT-LSTM-CRF model for sequence annotation. For the input text , first use the pre-trained language model BERT to extract the initial sequence feature information of the text data:

[0088]

[0089] where, is the initial sequence feature information, is the length of the text sequence, is the dimension of the BERT word vector feature, is the text data in the multimodal data.

[0090] In some embodiments, the pre-trained sequence feature extraction model can be an LSTM model. The extraction device can input the initial sequence feature information extracted by the BERT model into the LSTM model, utilize the sequence modeling ability of the LSTM to further enhance the context correlation information in the text sequence features, improve the performance of the model for complex sequence annotation tasks, and obtain text sequence feature information:

[0091] ,

[0092] where, is the text sequence feature information, is the length of the text sequence, is the size of the hidden layer dimension in the pre-trained sequence feature extraction model.

[0093] Furthermore, for the accuracy of text sequence annotation, the above step S103 may include:

[0094] Step S1031: Perform label annotation based on the text sequence feature information to obtain multiple text labels;

[0095] Step S1032: Analyze the dependency relationship between each text label and generate a predicted label sequence based on the dependency relationship;

[0096] Step S1033: Calculate the sequence scores of each predicted label sequence through a scoring function;

[0097] Step S1034: Perform normalization processing on the predicted label sequence based on the sequence scores to obtain sequence probability information;

[0098] Step S1035: Determine the text label sequence from the predicted label sequence according to the sequence probability information.

[0099] It should be noted that the sequence scores include emission scores and transition scores. The emission scores represent the label scores of each element in the predicted label sequence, and the transition scores represent the scores for transitioning from one label to another.

[0100] In some embodiments, the extraction device inputs the text sequence features into the CRF layer to obtain the BIO label sequence, and decodes to extract the events and their corresponding arguments in the text:

[0101]

[0102] Among them, represents the predicted label sequence, represents the length of the text sequence, represents the set of all possible label sequences, represents a possible label sequence, represents the scoring function, is the text sequence feature information, is the text data in the multimodal data, represents the sequence probability information, and the sequence probability information includes the predicted label sequence and the credible probability.

[0103] It should be noted that the above scoring function is used to calculate the scores of the predicted label sequences and consists of two parts: emission scores and transition scores. Among them, the emission scores refer to the label scores of each element in the sequence and are calculated from the probability distribution of the output labels. The transition scores refer to the scores for transitioning from one label to another. The CRF layer can effectively learn the constraint relationships between labels and improve the accuracy of the sequence labeling task.

[0104] In some embodiments, the loss function in the training process of the BERT-LSTM-CRF model is the negative log-likelihood function, as shown in the following formula:

[0105]

[0106] Among them, is the actual label sequence.

[0107] Step S20: Perform object detection on the image data in the multimodal data to obtain an object detection result.

[0108] It should be noted that in the multimodal event extraction task, due to the wide variety of argument entities participating in the event, the entity types in the image are often difficult to fully define in advance. Traditional object detection methods usually require predefining the categories of target entities and constructing a dataset based on the categories to train the model. This approach limits the detected entity categories and is difficult to be used in the multimodal event extraction task. Therefore, in this embodiment, the traditional object detection model can be improved from the perspective of the training strategy to make it an object detection model that does not depend on predefined categories. The improved object detection model can be trained only with bounding box annotations. At the same time, due to the lack of the limitation of the classification task, it also has the ability to detect new targets outside the training set.

[0109] It should be noted that the above object detection result includes the target bounding boxes in the image data and the judgment results on whether there are visual entities in each target bounding box.

[0110] In some embodiments, the extraction device can perform object detection on the image data through an object detection model (such as the YOLO model). The object detection can include object localization and object classification. The entity in the image data is labeled with a target bounding box through object localization, and whether there is a visual entity in each target bounding box is judged through object classification to obtain an object detection result, so that specific classification of the target is not required, effectively improving the object detection efficiency and ensuring the detection accuracy.

[0111] Further, referring to Figure 4 , Figure 4 which is a schematic diagram of object detection in an embodiment of the present invention. In order to accurately detect the visual entities in the image data, the above step S20 may include:

[0112] Step S201: Perform object localization on the image data in the multimodal data to obtain the entity position information in the image data;

[0113] Step S202: Label target bounding boxes in the image data based on the entity position information;

[0114] Step S203: Input the labeled image data into a pre-trained object detection model for object detection to obtain an object detection result.

[0115] In some embodiments, the extraction device can improve the YOLOv5 model from the perspective of the training strategy to make it an object detection model that does not depend on predefined categories. This model can be trained only with bounding box annotations. At the same time, due to the lack of the limitation of the classification task, it also has the ability to detect new targets outside the training set.

[0116] It should be noted that the training of the pre-trained object detection model includes two main tasks: the localization task is responsible for obtaining the position of the target entity in the image, and the classification task is responsible for identifying the category of the target entity. This article mainly improves the classification task, unifying all target entity categories into one category, thus transforming the original multi-class classification problem into a binary classification problem, that is, judging whether there is a target entity in the bounding box. The training loss of the model consists of three parts. The training loss function of the pre-trained object detection model includes:

[0117]

[0118] Among them, represents the training loss function of the pre-trained object detection model, represents the bounding box loss function, which is used to measure the difference between the predicted bounding box and the ground truth bounding box, represents the confidence loss function, which is used to measure the confidence that the object is contained in the target bounding box, is the binary classification loss function, which is used to judge whether there is a visual entity in the target bounding box.

[0119] Step S30: Obtain the text-image fine-grained information of the multi-modal data based on the text label sequence and the object detection result, and perform entity extraction on the multi-modal data based on the text-image fine-grained information to generate a text entity set and a visual entity set.

[0120] It should be noted that this embodiment can overcome the deficiency of the CLIP model in fine-grained concept understanding based on text event extraction and visual entity extraction, and achieve more accurate multi-modal information fusion. For the input text data and image data , first obtain a series of text entities and visual entities through text event extraction and visual entity extraction, so as to obtain the text entity set and the visual entity set .

[0121] In some embodiments, for the text event extraction model, the extraction device may adopt the RoBERTa model as the text encoder. The optimizer uses AdamW, and the learning rate is set to 2e-5. To enhance the learning ability of the CRF layer, a larger differential learning rate of 2e-3 is used. To avoid overfitting, an early stopping strategy is adopted to control the training of the model. When the performance on the validation set has not improved within 10 epochs, the training will stop, and the model with the best performance on the validation set is selected as the final model. The upper limit of the number of training epochs is 100, the batch size is 32, and the Dropout layer is 0.1. To ensure coverage of all texts, the maximum length of the input text is set to 256.

[0122] In some embodiments, for the visual entity extraction model, YOLOv5m is adopted as the object detector in this paper. The size of the input image is 640. The optimizer uses SGD, the learning rate is set to 1e-2, the batch size is 32, and the model with the best performance on the validation set is selected as the final model.

[0123] Step S40: Perform entity matching on the text entity set and the visual entity set to obtain an entity matching result.

[0124] It can be understood that in this embodiment, the text feature vectors of each text entity and the visual feature vectors of each visual entity in the text entity set and the visual entity set can be obtained by performing feature extraction on them. Entity matching is performed based on the text feature vectors and the visual feature vectors to obtain an entity matching result.

[0125] Further, referring to Figure 5 , Figure 5 which is a schematic diagram of the entity matching process in an embodiment of the present invention. To accurately match text entities and visual entities, the above step S40 may include:

[0126] Step S401: Input the text entity set and the visual entity set into a pre-trained feature encoding model for feature encoding to obtain the text feature representation information corresponding to the text entity set and the visual feature representation information corresponding to the visual entity set;

[0127] Step S402: Calculate the cosine similarity between each entity in the text entity set and the visual entity set based on the text feature representation information and the visual feature representation information;

[0128] Step S403: Perform entity matching according to the cosine similarity to obtain an entity matching result.

[0129] In some embodiments, the pre-trained feature encoding model may be a CLIP model. The extraction device inputs the extracted text entities and visual entities into the text encoder and visual encoder of CLIP respectively to obtain corresponding feature representations:

[0130]

[0131] where is the text feature representation information of the text entity set, is the visual feature representation information of the visual entity set, is the text encoder in the pre-trained feature encoding model, is the visual encoder in the pre-trained feature encoding model, is the text entity set, is the visual entity set, is the size of the hidden layer dimension in the pre-trained feature encoding model, represents the norm of the feature vector; represents the dot product of feature vectors.

[0132] In some embodiments, the extraction device can match the text entity and the visual entity pointing to the same argument role by calculating the cosine similarity between the text entity and the visual entity features. The calculation of the cosine similarity refers to the following formula:

[0133]

[0134] where is the text feature representation information, is the visual feature representation information, is the text entity and is the cosine similarity between the visual entity represents the dot product of feature vectors, represents the norm of the feature vector.

[0135] Meanwhile, to maximize the graph-text alignment ability of the CLIP model, in some embodiments, the extraction device also adopts a prompt integration technique to enhance the robustness of the inference result. Specifically, for each text entity, multiple different prompt templates are used for matching to obtain multiple visual entity matching results. Finally, the optimal matching result is determined through a voting mechanism.

[0136] Step S50: Based on the entity matching result, perform fine-grained graph-text alignment on the multimodal data, and input the multimodal data after fine-grained graph-text alignment into a pre-constructed event extraction model for event extraction.

[0137] It should be noted that the event extraction model can be a CLIP model. The CLIP model is good at extracting the global semantic information of text and images, but is relatively weak in understanding local semantic information. Since event information is usually complex and requires comprehensive understanding of global and local semantics, it is difficult to directly apply the CLIP model to complex semantic understanding scenarios such as multi-modal event extraction. Therefore, in this embodiment, by performing single-modal extraction on the text data and image data in the multi-modal data, then performing multi-modal fusion, performing fine-grained text-image alignment on the multi-modal data, and inputting the multi-modal data after fine-grained text-image alignment into the event extraction model for event extraction, the problem that the CLIP model cannot understand complex event semantics is effectively solved, the accuracy of event extraction is improved, and potential valuable events are effectively mined.

[0138] It can be understood that in this embodiment, by separately performing single-modal information extraction on the text data and image data in the multi-modal data, the text entities in the text data and the visual entities in the image data are obtained, and then text-image alignment is performed based on fine-grained information. By performing multi-modal fusion on the text entities and visual entities, event information can be accurately extracted from the multi-modal data.

[0139] In some embodiments, the extraction device can jointly extract event information from the text data and image data of the multi-modal data by analyzing the multi-modal data, respectively model the event information in the text and image using abstract semantic representation and visual scene graph, and then use the text-image pair data for weakly supervised alignment to achieve the joint extraction of text-image event information, so as to be able to extract richer event information and deeply mine the potential event value in the text-image data.

[0140] In some embodiments, the extraction device can use the Taiyi model as a multi-modal pre-training model for multi-modal information fusion. This model is based on the CLIP structure and is trained on a Chinese dataset to adapt to Chinese characteristics.

[0141] Further, referring to Figure 6 , Figure 6 is a schematic flowchart of model optimization in an embodiment of the present invention. In order to effectively improve the performance of the event extraction model, after the above step S50, the following steps may be included:

[0142] Step S501: Obtain the event extraction result output by the event extraction model;

[0143] Step S502: Perform performance evaluation on the event extraction model based on the predicted event type and predicted argument information to obtain a performance evaluation;

[0144] Step S503: Optimize the event extraction model according to the performance evaluation result.

[0145] It should be noted that the event extraction results include the predicted event types of the multimodal data and the predicted argument information, and the predicted argument information includes argument roles, text entities, and visual entities.

[0146] It should be explained that in this paper, precision (P), recall (R), and F1 value (F1) are used as the main indicators to evaluate the model performance. When the predicted event types, argument roles, text entities, and visual entities are consistent with the annotations, they are regarded as correct. Among them, in some embodiments, the intersection over union (IoU) for visual argument consistency can be set to be greater than 0.8. The performance evaluation results include:

[0147]

[0148]

[0149]

[0150] Among them, TP represents the number of correctly predicted entities, FP represents the number of incorrectly predicted entities, and FN represents the number of actual entities that the event extraction model fails to recognize. represents the model precision. represents the model recall. represents the harmonic mean of the model precision and the model recall.

[0151] In some embodiments, the extraction device can use an event element extraction dataset of open-source multimodal data to evaluate and optimize the event extraction model. For example, this dataset covers a variety of different event types: detection, maintenance, protection, configuration, and movement events, as well as a variety of argument roles: initiator, recipient, configuration device, time, and location. Each piece of data consists of a news text in the field to which the data to be analyzed belongs and its corresponding picture. The dataset includes a training set, a validation set, and a test set. For example, the training set contains 1136 pieces of data, the validation set contains 264 pieces of data, and the test set contains 200 pieces of data.

[0152] In this embodiment, feature extraction is performed on the text data in the multimodal data, and a text label sequence is generated based on the text feature extraction result. The text label sequence includes multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event. Object detection is performed on the image data in the multimodal data to obtain an object detection result. The object detection result includes the target bounding boxes in the image data and the judgment results on whether there are visual entities in each target bounding box. The fine-grained text-image information of the multimodal data is obtained based on the text label sequence and the object detection result, and entity extraction is performed on the multimodal data based on the fine-grained text-image information to generate a text entity set and a visual entity set. The text entity set and the visual entity set are subjected to entity matching to obtain an entity matching result. The entity matching result includes the matching results between the text entities and the visual entities belonging to the same argument role. Fine-grained text-image alignment is performed on the multimodal data based on the entity matching result, and the multimodal data after fine-grained text-image alignment is input into a pre-constructed event extraction model for event extraction. Since this embodiment obtains the fine-grained text-image information of the multimodal data, respectively extracts the text entity set and the visual entity set based on the fine-grained text-image information, and realizes fine-grained text-image alignment based on entity matching, it effectively analyzes the global and local semantic information, effectively improves the efficiency of multimodal event extraction, and thus can accurately extract events from multimodal data with complex semantic scenarios, effectively avoid missing potential event values, and realize accurately extracting the event information in the multimodal data.

[0153] In addition, an embodiment of the present invention also proposes a computer-readable storage medium, on which a multimodal event extraction program is stored. When the multimodal event extraction program is executed by a processor, the steps of the multimodal event extraction method described above are implemented.

[0154] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0155] The above computer-readable storage medium can be included in the multi-modal event extraction device; it can also exist independently without being assembled into the multi-modal event extraction device.

[0156] In addition, an embodiment of the present invention also proposes a computer program product, including a multi-modal event extraction program, and when the multi-modal event extraction program is executed by a processor, it implements the steps of the multi-modal event extraction method described above.

[0157] The specific implementation manner of the computer program product of the present invention is basically the same as that of each embodiment of the above multi-modal event extraction method, and will not be elaborated here.

[0158] Refer to Figure 7 , Figure 7 which is the structural block diagram of the first embodiment of the multi-modal event extraction device of the present invention.

[0159] As Figure 7 shown, the multi-modal event extraction device proposed by the embodiment of the present invention includes:

[0160] A text event analysis module 10, which is used to extract features from the text data in the multi-modal data and generate a text label sequence based on the text feature extraction result. The text label sequence contains multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event;

[0161] An image event analysis module 20, which is used to perform object detection on the image data in the multimodal data to obtain an object detection result, where the object detection result includes the object bounding boxes in the image data and the judgment result on whether there is a visual entity in each object bounding box;

[0162] A fine-grained analysis module 30, which is used to obtain the graphic and text fine-grained information of the multimodal data based on the text label sequence and the object detection result, and perform entity extraction on the multimodal data based on the graphic and text fine-grained information to generate a text entity set and a visual entity set;

[0163] An entity matching module 40, which is used to perform entity matching between the text entity set and the visual entity set to obtain an entity matching result, where the entity matching result includes the matching result between the text entity and the visual entity belonging to the same argument role;

[0164] An event extraction module 50, which is used to perform fine-grained graphic and text alignment on the multimodal data based on the entity matching result, and input the multimodal data after fine-grained graphic and text alignment into a pre-constructed event extraction model for event extraction.

[0165] Furthermore, the text feature extraction result includes text sequence feature information; the text event analysis module 10 is further used to input the text data in the multimodal data into a pre-trained language model for feature extraction to obtain the initial sequence feature information of the text data:

[0166]

[0167] Among them, is the initial sequence feature information, is the text sequence length, is the dimension of the word vector feature, is the text data in the multimodal data;

[0168] Input the initial sequence feature information into a pre-trained sequence feature extraction model to obtain text sequence feature information:

[0169] ,

[0170] Among them, is the text sequence feature information, is the text sequence length, is the size of the hidden layer dimension in the pre-trained sequence feature extraction model;

[0171] Generate a text label sequence based on the text sequence feature information.

[0172] Furthermore, the text event analysis module 10 is further configured to perform label annotation based on the text sequence feature information to obtain multiple text labels;

[0173] Analyze the dependency relationships between the text labels, and generate a predicted label sequence based on the dependency relationships;

[0174] Calculate the sequence scores of the predicted label sequences through a scoring function, where the sequence scores include emission scores and transition scores, the emission scores represent the label scores of the elements in the predicted label sequence, and the transition scores represent the scores for transitioning from one label to another;

[0175] Perform normalization processing on the predicted label sequences based on the sequence scores to obtain sequence probability information:

[0176]

[0177] Among them, represents the predicted label sequence, represents the length of the text sequence, represents the set of all possible label sequences, represents a possible label sequence, represents the scoring function, is the text sequence feature information, is the text data in the multimodal data, represents the sequence probability information, and the sequence probability information includes the predicted label sequence 's credible probability;

[0178] Determine the text label sequence from the predicted label sequences according to the sequence probability information.

[0179] Furthermore, the image event analysis module 20 is further configured to perform target localization on the image data in the multimodal data to obtain the entity position information in the image data;

[0180] Annotate the target bounding box in the image data based on the entity position information;

[0181] Input the annotated image data into a pre-trained object detection model for object detection to obtain object detection results, and the training loss function of the pre-trained object detection model includes:

[0182]

[0183] Among them, represents the training loss function of the pre-trained object detection model, Represents the bounding box loss function, which is used to measure the difference between the predicted bounding box and the ground truth bounding box. Represents the confidence loss function, which is used to measure the confidence that the object is contained in the target bounding box. Is the binary classification loss function, which is used to determine whether there is a visual entity in the target bounding box.

[0184] Furthermore, the entity matching module 40 is also used to input the text entity set and the visual entity set into the pre-trained feature encoding model for feature encoding, to obtain the text feature representation information corresponding to the text entity set and the visual feature representation information corresponding to the visual entity set:

[0185]

[0186] Among them, Is the text feature representation information of the text entity set, Is the visual feature representation information of the visual entity set, Is the size of the hidden layer dimension in the pre-trained feature encoding model;

[0187] Calculate the cosine similarity between each entity in the text entity set and the visual entity set based on the text feature representation information and the visual feature representation information:

[0188]

[0189] Among them, Is the text feature representation information, Is the visual feature representation information, Is the text entity And the visual entity The cosine similarity between them, Represents the dot product of feature vectors, Represents the norm of the feature vector;

[0190] Perform entity matching according to the cosine similarity to obtain the entity matching result.

[0191] Furthermore, the event extraction module 50 is also used to obtain the event extraction result output by the event extraction model. The event extraction result includes the predicted event type and predicted argument information of the multimodal data. The predicted argument information includes argument roles, text entities, and visual entities;

[0192] Based on the predicted event type and predicted argument information, perform performance evaluation on the event extraction model to obtain the performance evaluation result. The performance evaluation result includes:

[0193]

[0194]

[0195]

[0196] Among them, TP represents the number of correctly predicted entities, FP represents the number of incorrectly predicted entities, and FN represents the number of actual entities that the event extraction model fails to recognize. represents the precision of the model. represents the recall rate of the model. represents the harmonic mean of the precision and recall rate of the model.

[0197] Optimize the event extraction model according to the performance evaluation results.

[0198] In this embodiment, feature extraction is performed on the text data in the multimodal data, and a text label sequence is generated based on the text feature extraction results. The text label sequence includes multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event. Object detection is performed on the image data in the multimodal data to obtain an object detection result. The object detection result includes the target bounding boxes in the image data and the judgment results on whether there are visual entities in each target bounding box. The fine-grained text and image information of the multimodal data is obtained based on the text label sequence and the object detection result, and entity extraction is performed on the multimodal data based on the fine-grained text and image information to generate a text entity set and a visual entity set. The text entity set and the visual entity set are entity-matched to obtain an entity matching result. The entity matching result includes the matching results between the text entities and the visual entities belonging to the same argument role. Fine-grained text and image alignment is performed on the multimodal data based on the entity matching result, and the multimodal data after fine-grained text and image alignment is input into a pre-constructed event extraction model for event extraction. Since this embodiment obtains the fine-grained text and image information of the multimodal data, respectively extracts the text entity set and the visual entity set based on the fine-grained text and image information, and realizes fine-grained text and image alignment based on entity matching, effectively analyzing the global and local semantic information, effectively improving the efficiency of multimodal event extraction, so that events can be accurately extracted from multimodal data with complex semantic scenarios, effectively avoiding missing potential event values, and accurately extracting the event information in the multimodal data.

[0199] The multi-modal event extraction device provided by this application adopts the multi-modal event extraction method in the above-mentioned embodiment and can solve the technical problems of multi-modal event extraction. Compared with the prior art, the beneficial effects of the multi-modal event extraction device provided by this application are the same as those of the multi-modal event extraction method provided by the above-mentioned embodiment, and other technical features in the multi-modal event extraction device are the same as the features disclosed in the method of the above-mentioned embodiment, which will not be elaborated here.

[0200] It should be understood that the above is only an example for illustration and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not limit this.

[0201] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and there is no limitation here.

[0202] In addition, for the technical details not described in detail in this embodiment, reference can be made to the multi-modal event extraction method provided in any embodiment of the present invention, which will not be elaborated here.

[0203] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including the element.

[0204] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk), including several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0206] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A multimodal event extraction method, characterized in that, The multi-modal event extraction method includes: Extract features from the text data in the multi-modal data, and generate a text label sequence based on the text feature extraction result. The text label sequence includes multiple text labels, and the text labels include the text event types of the text events in the text data and the text entities corresponding to each text event; Perform object detection on the image data in the multi-modal data to obtain an object detection result. The object detection result includes the object bounding boxes in the image data and the judgment results on whether there are visual entities in each object bounding box; Obtain the fine-grained text-image information of the multi-modal data based on the text label sequence and the object detection result, and perform entity extraction on the multi-modal data based on the fine-grained text-image information to generate a text entity set and a visual entity set; Match the text entity set with the visual entity set to obtain an entity matching result. The entity matching result includes the matching results between the text entities and the visual entities belonging to the same argument role; Perform fine-grained text-image alignment on the multi-modal data based on the entity matching result, and input the multi-modal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction; The performing object detection on the image data in the multi-modal data to obtain an object detection result includes: Perform object localization on the image data in the multi-modal data to obtain the entity position information in the image data; Mark the object bounding boxes in the image data based on the entity position information; Input the marked image data into a pre-trained object detection model for object detection to obtain an object detection result. The training loss function of the pre-trained object detection model includes: Among them, represents the training loss function of the pre-trained object detection model, represents the bounding box loss function, and the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box, represents the confidence loss function, and the confidence loss function is used to measure the confidence that the object is contained in the target bounding box, is the binary classification loss function, and the binary classification loss function is used to determine whether there is a visual entity in the target bounding box; Before performing fine-grained text-image alignment on the multi-modal data based on the entity matching result and inputting the multi-modal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction, it further includes: Analyze the text data and image data of the multi-modal data to obtain the abstract semantic representation information in the text data and the visual scene graph in the image data; Model the text events and image events based on the abstract semantic representation information and the visual scene graph to obtain an initial extraction model; Perform weak supervision alignment on the initial extraction model according to the text-image pair data to obtain an event extraction model.

2. The multimodal event extraction method according to claim 1, wherein The text feature extraction result includes text sequence feature information. The extracting features from the text data in the multi-modal data and generating a text label sequence based on the text feature extraction result includes: Input the text data in the multi-modal data into a pre-trained language model for feature extraction to obtain the initial sequence feature information of the text data: Among them, is the initial sequence feature information, is the length of the text sequence, is the dimension of the word vector feature, is the text data in the multimodal data; Input the initial sequence feature information into a pre-trained sequence feature extraction model to obtain text sequence feature information: , Among them, is the text sequence feature information, is the length of the text sequence, is the size of the hidden layer dimension in the pre-trained sequence feature extraction model; Generate a text label sequence based on the text sequence feature information.

3. The multimodal event extraction method according to claim 2, wherein The generating a text label sequence based on the text sequence feature information includes: Perform label annotation based on the text sequence feature information to obtain multiple text labels; Analyze the dependencies between text tags and generate a predicted tag sequence based on the dependencies; Calculate the sequence scores of each predicted tag sequence through a scoring function, where the sequence scores include emission scores and transition scores. The emission scores represent the tag scores of each element in the predicted tag sequence, and the transition scores represent the scores for transitioning from one tag to another; Perform normalization processing on the predicted tag sequence based on the sequence scores to obtain sequence probability information: Among them, represents the predicted label sequence, represents the length of the text sequence, represents the set of all possible label sequences, represents a possible label sequence, represents the scoring function, is the feature information of the text sequence, is the text data in the multimodal data, represents the sequence probability information, and the sequence probability information includes the predicted label sequence 's credible probability; Determine the text tag sequence from the predicted tag sequence according to the sequence probability information.

4. The multimodal event extraction method according to claim 1, wherein The entity matching of the text entity set and the visual entity set to obtain an entity matching result includes: Input the text entity set and the visual entity set into a pre-trained feature encoding model for feature encoding to obtain the text feature representation information corresponding to the text entity set and the visual feature representation information corresponding to the visual entity set: Among them, is the text feature representation information of the text entity set, is the visual feature representation information of the visual entity set, is the text encoder in the pre-trained feature encoding model, is the visual encoder in the pre-trained feature encoding model, is the text entity set, is the visual entity set, is the size of the hidden layer dimension in the pre-trained feature encoding model; Calculate the cosine similarity between each entity in the text entity set and the visual entity set based on the text feature representation information and the visual feature representation information: Among them, is the text feature representation information, is the visual feature representation information, is the text entity and the visual entity the cosine similarity between them, represents the dot product of feature vectors, represents the norm of the feature vector; Perform entity matching according to the cosine similarity to obtain an entity matching result.

5. The multimodal event extraction method according to any one of claims 1 to 4, characterized in that After the fine-grained text-image alignment of the multimodal data based on the entity matching result and inputting the fine-grained text-image aligned multimodal data into a pre-constructed event extraction model for event extraction, it further includes: Obtain the event extraction result output by the event extraction model. The event extraction result includes the predicted event type of the multimodal data and the predicted argument information. The predicted argument information includes argument roles, text entities, and visual entities; Perform performance evaluation on the event extraction model based on the predicted event type and predicted argument information to obtain a performance evaluation result. The performance evaluation result includes: Among them, TP represents the number of correctly predicted entities, FP represents the number of incorrectly predicted entities, and FN represents the number of actual entities that the event extraction model fails to recognize. represents the precision of the model, represents the recall of the model, represents the harmonic mean of the precision and recall of the model; Optimize the event extraction model according to the performance evaluation result.

6. A multimodal event extraction device, characterized in that, The multimodal event extraction device includes: A text event analysis module for extracting features from the text data in the multimodal data and generating a text tag sequence based on the text feature extraction result. The text tag sequence contains multiple text tags, and the text tags include the text event types of the text events in the text data and the text entities corresponding to each text event; An image event analysis module for performing object detection on the image data in the multimodal data to obtain an object detection result. The object detection result includes the target bounding boxes in the image data and the judgment results on whether there are visual entities in each target bounding box; A fine-grained analysis module for obtaining the text-image fine-grained information of the multimodal data based on the text tag sequence and the object detection result, and performing entity extraction on the multimodal data based on the text-image fine-grained information to generate a text entity set and a visual entity set; An entity matching module for performing entity matching on the text entity set and the visual entity set to obtain an entity matching result. The entity matching result includes the matching results between the text entities and the visual entities belonging to the same argument role; An event extraction module, configured to perform fine-grained text-image alignment on the multimodal data based on the entity matching result, and input the multimodal data after fine-grained text-image alignment into a pre-constructed event extraction model for event extraction; The image event analysis module is further configured to perform object localization on the image data in the multimodal data to obtain entity position information in the image data; annotate a target bounding box in the image data based on the entity position information; input the annotated image data into a pre-trained object detection model for object detection to obtain an object detection result, and the training loss function of the pre-trained object detection model includes: Among them, represents the training loss function of the pre-trained object detection model, represents the bounding box loss function, and the bounding box loss function is used to measure the difference between the predicted bounding box and the true bounding box. represents the confidence loss function, and the confidence loss function is used to measure the confidence that the object is contained in the target bounding box. is the binary classification loss function, and the binary classification loss function is used to determine whether there is a visual entity in the target bounding box; The event extraction module is further configured to analyze the text data and image data of the multimodal data to obtain abstract semantic representation information in the text data and a visual scene graph in the image data; model text events and image events based on the abstract semantic representation information and the visual scene graph to obtain an initial extraction model; perform weak supervision alignment on the initial extraction model according to text-image pair data to obtain an event extraction model.

7. A multimodal event extraction device, characterized in that, The multimodal event extraction device includes: a memory, a processor, and a multimodal event extraction program stored on the memory and executable on the processor, and the multimodal event extraction program is configured to implement the multimodal event extraction method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a multimodal event extraction program, and when the multimodal event extraction program is executed by a processor, it implements the multimodal event extraction method according to any one of claims 1 to 5.

9. A computer program product, characterized in that, The computer program product includes a multimodal event extraction program, and when the multimodal event extraction program is executed by a processor, it implements the steps of the multimodal event extraction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Structured information extraction method and device based on multi-element labeling strategy

    CN113836891A

  • Multi-modal element extraction method, system and device

    CN118410163A