Event extraction model processing method and device, equipment and storage medium

Through cross-modal reconstruction and loss optimization event extraction model for multimodal data, the inconsistency problem caused by modal independent training is solved, and the accuracy and consistency of event extraction is improved.

CN120338084APending Publication Date: 2025-07-18TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510452161.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

After independently training each mode, the existing multimodal data event extraction model leads to inconsistent extraction results of different modes, affecting the accuracy of integrated event extraction.

Method used

Through the event extraction model, the multimodal data is performed on the event extraction process of each mode, and the cross-modal data reconstruction loss and event extraction loss are determined. The model is optimized based on these losses to obtain the target event extraction model.

Benefits of technology

It improves the accuracy and consistency of extraction between different modes, reduces the inconsistency problems caused by independent mode training, and improves the event extraction effect of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338084A_ABST
    Figure CN120338084A_ABST
Patent Text Reader

Abstract

The invention relates to an event extraction model processing method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: performing event extraction processing of each mode on first multi-mode data through an event extraction model to obtain first event information of each mode; performing cross-modal data reconstruction based on the first event information of each modal to obtain cross-modal reconstruction data of each modal; determining the cross-modal reconstruction loss of each modal based on the first multi-modal data and the cross-modal reconstruction data of each modal, and determining the first event extraction loss of each modal based on the first event information of each modal and the corresponding event tag; and according to the cross-modal reconstruction loss and the first event extraction loss, performing parameter optimization on the event extraction model to obtain a target event extraction model. By adopting the target event extraction model obtained by the method, the accuracy of an event extraction result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, device, and storage medium for processing an event extraction model. Background Art

[0002] In the real world, a large amount of data exists in unstructured form, such as news reports, social media, image descriptions, video commentaries, etc. By performing event extraction on this data to extract structured event information, key support can be provided for numerous intelligent systems based on the extracted structured event information.

[0003] In recent years, with the rapid development of computer technology, more and more data presents multi-modal characteristics. Multi-modal data refers to the combination of various different forms such as text, vision, and audio. Currently, event extraction for multi-modal data usually extracts data of corresponding modalities through event extraction models specified for each modality, and then integrates the extraction results of different modalities. However, event extraction models are often independently trained only based on the training data of the corresponding modality during training. The extraction results of different-modal event extraction models obtained in this independent training manner may be inconsistent, resulting in poor accuracy of the event extraction results obtained by integrating the extraction results of different modalities. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, device, and storage medium for processing an event extraction model that can improve the accuracy of event extraction results for the above technical problems.

[0005] In a first aspect, this application provides a method for processing an event extraction model. The method includes:

[0006] Performing event extraction processing on the first multi-modal data for each modality through the event extraction model to obtain first event information for each modality;

[0007] Performing cross-modal data reconstruction based on the first event information for each modality to obtain cross-modal reconstruction data for each modality;

[0008] Determining cross-modal reconstruction losses for each modality based on the first multi-modal data and the cross-modal reconstruction data for each modality, and determining first event extraction losses for each modality based on the first event information for each modality and corresponding event labels;

[0009] Optimizing the parameters of the event extraction model according to the cross-modal reconstruction losses and the first event extraction losses to obtain a target event extraction model.

[0010] In a second aspect, the present application also provides a processing device for an event extraction model. The device includes:

[0011] An event extraction module, configured to perform event extraction processing on each modality of the first multi-modal data through the event extraction model to obtain first event information for each modality;

[0012] A data reconstruction module, configured to perform cross-modal data reconstruction based on the first event information of each modality to obtain cross-modal reconstruction data for each modality;

[0013] A loss determination module, configured to determine the cross-modal reconstruction loss for each modality based on the first multi-modal data and the cross-modal reconstruction data for each modality, and determine the first event extraction loss for each modality based on the first event information of each modality and the corresponding event label;

[0014] A parameter optimization module, configured to optimize the parameters of the event extraction model according to the cross-modal reconstruction loss and the first event extraction loss to obtain a target event extraction model.

[0015] In one embodiment, the event extraction model includes event extraction branches for each modality, and the target event extraction model includes target event extraction branches for each modality; the event extraction module is further configured to:

[0016] Perform event extraction processing on each modality of the first multi-modal data through the event extraction branches for each modality to obtain first event information for each modality;

[0017] The parameter optimization module is further configured to:

[0018] Optimize the parameters of the event extraction branches for the corresponding modalities according to the cross-modal reconstruction loss and the first event extraction loss for each modality to obtain target event extraction branches for each modality.

[0019] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the device further includes:

[0020] A fusion module, configured to fuse the first event information of each modality through the fusion network to obtain first fused event information;

[0021] The loss determination module is further configured to determine a first fusion loss based on the first fused event information and the fused event label;

[0022] The parameter optimization module is further configured to optimize the parameters of the fusion network based on the first fusion loss to obtain a target fusion network.

[0023] In one embodiment, the fusion module is further configured to:

[0024] Obtain the information confidence levels respectively corresponding to the first event information of each of the modalities;

[0025] Based on the information confidence levels of the first event information of each of the modalities, fuse the first event information of each of the modalities through the fusion network to obtain first fused event information.

[0026] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network;

[0027] The event extraction module is further configured to perform event extraction processing for each modality on the second multi-modal data through the target event extraction branches of each of the modalities to obtain second event information of each of the modalities;

[0028] The apparatus further includes a fusion module, configured to fuse the second event information of each of the modalities through the fusion network to obtain second fused event information;

[0029] The loss determination module is further configured to determine a second fusion loss based on the second fused event information and the fusion event label;

[0030] The parameter optimization module is further configured to optimize the parameters of the fusion network based on the second fusion loss to obtain a target fusion network.

[0031] In one embodiment, the apparatus further includes a pre-training module, configured to:

[0032] Perform event extraction processing for each of the modalities on the third multi-modal data through an initial event extraction model to obtain third event information of each of the modalities;

[0033] Perform data reconstruction of the original modality based on the third event information of each of the modalities to obtain original modality reconstruction data of each of the modalities;

[0034] Optimize the parameters of the initial event extraction model based on the original modality reconstruction data of each of the modalities and the third event information of each of the modalities to obtain the event extraction model.

[0035] In one embodiment, the pre-training module is further configured to:

[0036] Determine the original modality reconstruction loss of each of the modalities based on the third multi-modal data and the original modality reconstruction data of each of the modalities, and determine the second event extraction loss of each of the modalities based on the third event information of each of the modalities and the corresponding event label;

[0037] Optimize the parameters of the initial event extraction model according to the original modal reconstruction loss and the second event extraction loss to obtain the event extraction model.

[0038] In one embodiment, the apparatus includes an inference module, configured to:

[0039] Obtain multi-modal data to be processed;

[0040] Perform event extraction processing on each modality of the multi-modal data to be processed through a target event extraction model to obtain event information of each modality of the multi-modal data to be processed;

[0041] Determine the event information of the multi-modal data to be processed based on the event information of each modality of the multi-modal data to be processed.

[0042] In one embodiment, the target event extraction model includes a target fusion network and target event extraction branches for each modality; the inference module is further configured to:

[0043] Obtain the information confidence corresponding to the event information of each modality;

[0044] Through the target fusion network, fuse the event information of each modality based on the information confidence of the event information of each modality to obtain event information.

[0045] In one embodiment, the inference module is further configured to:

[0046] Obtain the modality weights corresponding to each modality;

[0047] Fuse the event information of each modality based on the modality weights of each modality to obtain event information.

[0048] In one embodiment, the apparatus further includes an event retrieval module, configured to:

[0049] Add the event information of the multi-modal data to be processed to an event information library;

[0050] When receiving a retrieval request, extract a query condition from the retrieval request;

[0051] In the event information library, search for target event information that meets the query condition;

[0052] Return the multi-modal data corresponding to the target event information as a retrieval result to the initiator of the retrieval request.

[0053] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above methods are implemented.

[0054] In a fourth aspect, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above methods are implemented.

[0055] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of any one of the above methods are implemented.

[0056] For the above event extraction model processing method, device, computer device, storage medium, and computer program product, through the event extraction model, event extraction processing is performed on each modality of the first multi-modal data to obtain the first event information of each modality. Based on the first event information of each modality, cross-modal data reconstruction is performed to obtain the cross-modal reconstruction data of each modality. Based on the first multi-modal data and the cross-modal reconstruction data of each modality, the cross-modal reconstruction loss of each modality is determined. According to the cross-modal reconstruction loss, the parameters of the event extraction model are optimized, so that the model can capture the mutual relationship between different modalities, help the model understand the connection between different modalities, reduce the inconsistency between the results of independent modality extraction, and avoid the inconsistency problem caused by independent modality training. Based on the first event information of each modality and the corresponding event label, the first event extraction loss of each modality is determined. According to the first event extraction loss, the parameters of the event extraction model are optimized, which can improve the accuracy of the model for event extraction. Therefore, by jointly optimizing the cross-modal reconstruction loss and the event extraction loss, the model can simultaneously improve the extraction accuracy of different modalities and the consistency between modalities, thereby improving the event extraction effect of multi-modal data, avoiding the problem of error accumulation and accuracy decline that may occur during independent training, and further improving the accuracy of event information extraction when the target event extraction model can be used for event information extraction of multi-modal data subsequently. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is an application environment diagram of the processing method of the event extraction model in an embodiment;

[0058] Figure 2 It is a schematic flowchart of the processing method of the event extraction model in an embodiment;

[0059] Figure 3Schematic diagram for cross-modal reconstruction verification of text data in an embodiment;

[0060] Figure 4 Schematic diagram for cross-modal reconstruction verification of visual data in an embodiment;

[0061] Figure 5 Schematic diagram for cross-modal reconstruction verification of audio data in an embodiment;

[0062] Figure 6 Flow schematic diagram of steps for a pre-trained event extraction model in an embodiment;

[0063] Figure 7 Schematic diagram for original-modal reconstruction verification of text data in an embodiment;

[0064] Figure 8 Schematic diagram for original-modal reconstruction verification of visual data in an embodiment;

[0065] Figure 9 Schematic diagram for original-modal reconstruction verification of audio data in an embodiment;

[0066] Figure 10 Flow schematic diagram of steps for event information extraction in an embodiment;

[0067] Figure 11 Flow schematic diagram of a processing method for an event extraction model in another embodiment;

[0068] Figure 12 Flow schematic diagram of a processing method for an event extraction model in another embodiment;

[0069] Figure 13 Structure block diagram of a processing device for an event extraction model in an embodiment;

[0070] Figure 14 Structure block diagram of a processing device for an event extraction model in an embodiment;

[0071] Figure 15 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0072] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0073] The processing method for an event extraction model provided by an embodiment of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The processing method of the above event extraction model can be executed independently by the terminal 102 or the server 104, or can be executed through the interaction between the terminal 102 and the server 104. Taking the server 104 executing alone as an example, in one embodiment, the server performs event extraction processing on each modality of the first multi-modal data through the event extraction model to obtain the first event information of each modality; based on the first event information of each modality, cross-modal data reconstruction is performed to obtain the cross-modal reconstruction data of each modality; based on the first multi-modal data and the cross-modal reconstruction data of each modality, the cross-modal reconstruction loss of each modality is determined, and based on the first event information of each modality and the corresponding event label, the first event extraction loss of each modality is determined; according to the cross-modal reconstruction loss and the first event extraction loss, the parameters of the event extraction model are optimized to obtain the target event extraction model. For example, the first multi-modal data includes the original first modality data and the original second modality data, the event extraction model includes an event extraction branch of the first modality and an event extraction branch of the second modality, and the server 104 performs event extraction processing on the original first modality data and the original second modality data respectively through the event extraction branch of the first modality and the event extraction branch of the second modality to obtain the first event information of the first modality and the first event information of the second modality, and inputs the first event information of the first modality into the cross-modal reconstruction model 1 to process the first event information of the first modality through the cross-modal reconstruction model 1 to obtain the cross-modal reconstruction data of the second modality, inputs the first event information of the second modality into the cross-modal reconstruction model 2 to process the first event information of the second modality through the cross-modal reconstruction model 2 to obtain the cross-modal reconstruction data of the first modality, and based on the cross-modal reconstruction data of the second modality and the original first modality data, the cross-modal reconstruction loss of the first modality is determined, based on the first event information of the first modality and the event label of the first modality, the first event extraction loss of the first modality is determined, based on the cross-modal reconstruction data of the first modality and the original second modality data, the cross-modal reconstruction loss of the second modality is determined, based on the first event information of the second modality and the event label of the second modality, the first event extraction loss of the second modality is determined, based on the cross-modal reconstruction data of the first modality and the first event extraction loss of the first modality, the parameters of the event extraction branch of the first modality are optimized, based on the cross-modal reconstruction data of the second modality and the first event extraction loss of the second modality, the parameters of the event extraction branch of the second modality are optimized to obtain the target event extraction model.

[0074] It can be understood that after the server 104 obtains the trained target event extraction model, it can also send the target event extraction model to the terminal 102, so that the terminal 102 can use this model to perform event extraction on multi-modal data.

[0075] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 102 and the server 104 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions in this regard.

[0076] In one embodiment, as Figure 2 shown, a processing method for an event extraction model is provided. Taking the method applied to Figure 1 the computer device in

[0077] S202, through the event extraction model, perform event extraction processing on each modality of the first multi-modal data to obtain first event information for each modality.

[0078] Among them, the event extraction model is a model used to identify and extract event information from multi-modal data. Specifically, it can be a neural network model based on deep learning, such as a model based on a convolutional neural network (CNN), a model based on a recurrent neural network (RNN) and a long short-term memory network (LSTM), or any one of the models based on Transformer. Multi-modal data refers to a data set containing data from different modalities, where each modality represents a different type of information source or expression method. Multi-modal data can include at least two different modalities of data, and the modalities corresponding to the multi-modal data can specifically include at least two of the text modality, the audio modality, and the visual modality.

[0079] In the embodiments of the present application, the first multi-modal data may specifically be multi-modal data describing a first sample event, and is also sample data for training an event extraction model. The first multi-modal data is composed of multiple different modal data types, and these data types can describe the same event from different perspectives and dimensions, thereby providing more comprehensive and accurate event information. For example, a certain first sample event is a traffic accident, and the first multi-modal data corresponding to the first sample event may include text data 1, visual data 1, and audio data 1. Among them, the text data 1 is "On May 3, 2023, a serious car accident occurred in California, resulting in 3 deaths", the visual data 1 is "Photos of the traffic accident scene, showing the collided vehicles and damaged traffic facilities", and the audio data 1 is "Recordings of the accident scene obtained from nearby cameras or passers-by's mobile phones, including the sounds of vehicle collisions, conversations of passers-by or drivers, etc.".

[0080] Event Extraction refers to the process of identifying and extracting event information from data. The event information may specifically include event types and event elements. The event type refers to the category to which the event belongs, usually a label describing the essential characteristics of the event, such as typhoon, meeting, press conference, etc. The event element refers to the specific information or components related to the event type and is the key element describing an event. The event elements may include TriggerWord, Participants, Time, Location, Cause, Effect, etc. The TriggerWord is a verb or verb phrase indicating the occurrence of an event and is a key element in event extraction. For example, in "An explosion occurred", the trigger word is "occurred". The Participants refer to the entities, people, or things related to the event. The Time refers to the time when the event occurred. The Location refers to the location where the event occurred. The Cause refers to the reason or background that led to the event, usually the triggering factor of the event. The Effect refers to the result or consequence of the event, indicating the impact brought about after the event occurred. For example, in "The earthquake caused the building to collapse", the result is "the building collapsed".

[0081] The first event information of each modality refers to the event information extracted based on each modality data in the multi-modal data. For example, the first multi-modal data includes data of the first modality and data of the second modality. The first event information of the first modality is extracted based on the data of the first modality, and the first event information of the second modality is extracted based on the data of the second modality. Among them, the first modality may be at least one of the text modality, audio modality, and visual modality, and the second modality is a modality different from the first modality among the text modality, audio modality, and visual modality.

[0082] Specifically, the computer device can obtain the first multi-modal data from the training dataset and input the first multi-modal data into the event extraction model to be trained. Through the event extraction branches corresponding to each modality of the event extraction model, event extraction processing is respectively performed on the data of different modalities in the first multi-modal data to obtain the first event information corresponding to the data of each modality.

[0083] In one embodiment, the event extraction branches corresponding to each modality respectively include an encoder and a decoder. For the data of a certain modality, the data of this modality can be input into the event extraction branch of this modality. The encoder of the event extraction branch of this modality encodes the input data to obtain an encoded vector, and the encoded vector data is input into the decoder of the event extraction branch of this modality. The decoder decodes the input encoded vector to obtain the first event information corresponding to the data of this modality. By performing the above processing on the data of each modality, the first event information corresponding to the data of each modality can be obtained.

[0084] In one embodiment, the computer device can first perform feature processing on the first multi-modal data to obtain the features of each modality, and input the features of different modalities into the event extraction model to be trained. Through the event extraction branches corresponding to each modality of the event extraction model, event extraction processing is respectively performed on the features of each modality to obtain the first event information of each modality.

[0085] In one embodiment, the first multi-modal data includes text data, and the computer device performs feature processing on the text data to obtain text features.

[0086] Specifically, the computer device can perform word segmentation processing on the text data to obtain a word segmentation result, remove stop words from the word segmentation result to obtain each word, and input each word into the text feature extraction model to convert each word into a numerical text feature through the text feature extraction model. The text feature extraction model can be any one of Word2Vec, GloVe, or BERT.

[0087] In one embodiment, the first multi-modal data includes visual data, and the computer device performs feature processing on the visual data to obtain visual features.

[0088] Specifically, the computer device can input the visual data into the visual feature extraction model to perform feature extraction on the visual data through the visual feature extraction model to obtain visual features.

[0089] Among them, the visual feature extraction model can be a model based on a convolutional neural network (CNN). The CNN automatically learns the local features and global structure of an image through multiple layers of convolution and pooling operations, thereby generating a discriminative image feature vector. After being processed by the CNN, the original pixel matrix of the image is converted into a high-level feature vector, which can better capture the key information in the image.

[0090] In one embodiment, the first multi-modal data includes audio data, and the computer device processes the audio data to obtain audio features.

[0091] Among them, the audio features can be Mel Frequency Cepstral Coefficients (MFCC). MFCC is a feature widely used in audio processing. Based on the auditory characteristics of the human ear, the audio signal is converted into frequency-related coefficients, which can reflect the spectral envelope of the audio signal and are very effective for tasks such as speech recognition and music classification.

[0092] Specifically, the computer device can perform a Short-Time Fourier Transform (STFT) on the audio data to obtain the Fourier transform result, extract the Mel Frequency Cepstral Coefficients of the audio data based on the Fourier transform result, and use the extracted Mel Frequency Cepstral Coefficients as the audio features of the audio data.

[0093] Combined with the following examples, the above embodiments are described. For example, the first multi-modal data includes text data T, visual data I, and audio data S, and text features , visual features and audio features are obtained based on the text data T, visual data I, and audio data S. , visual features and audio features are respectively input into the event extraction branches of the multi-layer Transformer model based on the corresponding modalities. Assuming that the feature dimensions of the text features , visual features and audio features are the same, for any modality of features, they are characterized as feature sequences X=[ x 1 , x 2 ,⋯, x n ] , encode X into a vector C through the Encoder of the Transformer model, C = Transformer(X), then input the encoded vector C into the Decoder of the Transformer model, and decode the encoded vector C through the Decoder to obtain the output first event information Y. The first event information Y is represented in sequence form, Y = [E, e 1 , e 2 ,⋯, e n ] , where E is the event category, is the i-th event element, by performing the above processing on the text features , visual features and audio features respectively, thus obtaining the first event information corresponding to the text data T , the first event information corresponding to the visual data I and the first event information corresponding to the audio data S .

[0094] S204, perform cross-modal data reconstruction based on the first event information of each modality to obtain cross-modal reconstruction data of each modality.

[0095] Among them, cross-modal refers to the interaction and conversion between different types of data or information, such as converting data of the first modality into data of the second modality. Data reconstruction refers to the process of generating complete data or information related to it using a model or algorithm based on existing data or information.

[0096] The cross-modal data reconstruction in the embodiments of this application refers to reconstructing data of another modality based on data of one modality, such as generating image data or audio data based on text data. Specifically, it can be generating image data or audio data based on the first event information of the text modality.

[0097] Specifically, after the computer device obtains the first event information of each modality, for the first event information of the first modality, it can input it into the cross-modal reconstruction model corresponding to the second modality, and generate cross-modal reconstruction data of the second modality through the cross-modal reconstruction model corresponding to the second modality based on the first event information of the first modality; for the first event information of the second modality, it can input it into the cross-modal reconstruction model corresponding to the first modality, and generate cross-modal reconstruction data of the first modality through the cross-modal reconstruction model corresponding to the first modality based on the first event information of the second modality; by performing the above processing on the first event information of each modality, cross-modal reconstruction data of each modality can be obtained.

[0098] Among them, the first modality may be at least one of a text modality, an audio modality, and a visual modality, and the second modality is a modality different from the first modality among the text modality, the audio modality, and the visual modality.

[0099] For example, the first multi-modal data includes text data, visual data, and audio data. By performing event extraction processing on the multi-modal data, the first event information corresponding to the text data, the first event information corresponding to the visual data, and the first event information corresponding to the audio data can be obtained. Based on the first event information corresponding to the text data, cross-visual modality data reconstruction is performed to obtain cross-modal reconstructed visual data corresponding to the text data. Based on the first event information corresponding to the text data, cross-audio modality data reconstruction is performed to obtain cross-modal reconstructed audio data corresponding to the text data. Based on the first event information corresponding to the audio data, cross-visual modality data reconstruction is performed to obtain cross-modal reconstructed visual data corresponding to the audio data. Based on the first event information corresponding to the audio data, cross-text modality data reconstruction is performed to obtain cross-modal reconstructed text data corresponding to the audio data. Based on the first event information corresponding to the visual data, cross-text modality data reconstruction is performed to obtain cross-modal reconstructed text data corresponding to the visual data. Based on the first event information corresponding to the visual data, cross-audio modality data reconstruction is performed to obtain cross-modal reconstructed audio data corresponding to the visual data.

[0100] S206. Determine the cross-modal reconstruction loss of each modality based on the first multi-modal data and the cross-modal reconstructed data of each modality, and determine the first event extraction loss of each modality based on the first event information of each modality and the corresponding event label.

[0101] Among them, loss, also known as the loss value or the loss function value, is used to guide model training. The event label refers to the pre-annotated event information. For the text modality, the corresponding event label is the annotated event information of the text data. For the visual modality, the corresponding event label is the annotated event information of the visual data. For the audio modality, the corresponding event label is the annotated event information of the audio data.

[0102] Specifically, after the computer prepares to obtain the first event information of each modality, it can obtain the event labels corresponding to the data of each modality. For the data of the first modality, based on the first event information and event labels corresponding to the data of the first modality, it can determine the first event extraction loss corresponding to the data of the first modality, and based on the data of the first modality and the cross-modal reconstruction data of the second modality, determine the cross-modal reconstruction loss corresponding to the data of the first modality; for the data of the second modality, based on the first event information and event labels corresponding to the data of the second modality, it can determine the first event extraction loss corresponding to the data of the second modality, and based on the data of the second modality and the cross-modal reconstruction data of the first modality, determine the cross-modal reconstruction loss corresponding to the data of the second modality; by performing the above processing on the first event information and cross-modal reconstruction data of each modality, the first event extraction loss and cross-modal reconstruction loss of each modality can be obtained.

[0103] The above embodiments are described in conjunction with the following examples. For example, the first multi-modal data includes text data T, visual data I, and audio data S, as Figure 3 shown, the text data T corresponds to the first event information , event label , cross-modal reconstruction visual data , cross-modal reconstruction audio data , then based on the first event information , event label determine the first event extraction loss corresponding to the text data T, based on the cross-modal reconstruction visual data and the visual data I to determine the first cross-modal loss corresponding to the text data T, based on the cross-modal reconstruction audio data and the audio data S to determine the second cross-modal loss corresponding to the text data T, based on the first cross-modal loss and the second cross-modal loss , determine the cross-modal reconstruction loss corresponding to the text data T; as Figure 4 shown, the visual data I corresponds to the first event information , event label , cross-modal reconstruction text data , cross-modal reconstruction audio data , then based on the first event information , event label determine the first event extraction loss corresponding to the visual data I, based on the cross-modal reconstruction text data and the text data T to determine the first cross-modal loss , based on cross-modal reconstructed audio data and the audio data S to determine the second cross-modal loss corresponding to the visual data I , based on the first cross-modal loss and the second cross-modal loss , determine the cross-modal reconstruction loss corresponding to the visual data I ; as Figure 5 shown, the audio data S corresponds to the first event information , event label , cross-modal reconstructed text data , cross-modal reconstructed visual data , then based on the first event information , event label to determine the first event extraction loss corresponding to the audio data S , based on the cross-modal reconstructed text data and the text data T to determine the first cross-modal loss corresponding to the audio data S , based on the cross-modal reconstructed visual data and the visual data I to determine the second cross-modal loss corresponding to the audio data S , based on the first cross-modal loss and the second cross-modal loss , determine the cross-modal reconstruction loss corresponding to the audio data S .

[0104] S208, according to the cross-modal reconstruction loss and the first event extraction loss, optimize the parameters of the event extraction model to obtain the target event extraction model.

[0105] Specifically, after the computer device obtains the cross-modal reconstruction loss corresponding to the data of each modality and the first event extraction loss corresponding to the data of each modality, for the data of the first modality, based on the cross-modal reconstruction loss and the first event extraction loss corresponding to it, determine the training loss corresponding to the data of the first modality. For the data of the second modality, based on the cross-modal reconstruction loss and the first event extraction loss corresponding to it, determine the training loss corresponding to the data of the second modality; by performing the above processing on the cross-modal reconstruction loss and the first event extraction loss corresponding to the data of each modality, the training loss corresponding to the data of each modality is obtained, and the parameters of the event extraction model are adjusted based on the training loss corresponding to the data of each modality, so as to obtain the adjusted event extraction model, and then return to execute step S202 until the convergence condition is reached to obtain the trained target event extraction model.

[0106] Among them, the parameters of the event extraction model can specifically be at least one of the weights and biases in the event extraction model. The convergence condition refers to the condition in the model training process that, according to a pre-set standard or threshold, when the training loss value or other evaluation metrics change very slightly or reach a stable state in a continuous number of iterations, it is considered that the model parameters have been fully optimized, the model performance tends to be stable, and at this time the training process can be terminated. The convergence condition in the embodiments of the present application can be any one of the following: 1. The change range of the loss is low. For example, in a continuous number of iteration cycles (epochs), the decrease range of the total loss value is lower than the pre-set threshold; 2. The verification metrics are stable. For example, the performance metrics (such as accuracy, mean squared error, etc.) on the validation set remain stable in a continuous number of cycles without obvious improvement; 3. The maximum number of iterations is reached. For example, the training process runs to the pre-set maximum number of iterations or training rounds.

[0107] In the above method for processing the event extraction model, through the event extraction model, event extraction processing is performed on each modality of the first multi-modal data to obtain the first event information of each modality. Based on the first event information of each modality, cross-modal data reconstruction is performed to obtain the cross-modal reconstruction data of each modality. Based on the first multi-modal data and the cross-modal reconstruction data of each modality, the cross-modal reconstruction loss of each modality is determined. According to the cross-modal reconstruction loss, the parameters of the event extraction model are optimized, so that the model can capture the mutual relationship between different modalities, help the model understand the connection between different modalities, reduce the inconsistency between the results of independent modality extraction, and avoid the inconsistency problem caused by independent modality training; based on the first event information of each modality and the corresponding event labels, the first event extraction loss of each modality is determined. According to the first event extraction loss, the parameters of the event extraction model are optimized, which can improve the accuracy of the model for event extraction; therefore, by jointly optimizing the cross-modal reconstruction loss and the event extraction loss, the model can simultaneously improve the extraction accuracy of different modalities and the consistency between modalities, thereby improving the event extraction effect of multi-modal data, avoiding the problem of error accumulation and accuracy decline that may occur during independent training, and further improving the accuracy of event information extraction when the target event extraction model can be used for event information extraction of multi-modal data subsequently.

[0108] In one embodiment, the event extraction model includes event extraction branches for each modality, and the target event extraction model includes target event extraction branches for each modality; the process by which the computer device performs event extraction processing for each modality on the first multi-modal data through the event extraction model to obtain the first event information for each modality includes the following steps: performing event extraction processing for each modality on the first multi-modal data through the event extraction branches for each modality to obtain the first event information for each modality; the computer device optimizes the parameters of the event extraction branches for the corresponding modality based on the cross-modal reconstruction loss and the first event extraction loss for each modality to obtain the target event extraction branches for each modality.

[0109] Among them, the event extraction branch corresponding to each modality refers to an independent event extraction sub-model or sub-network structure designed and deployed respectively for data of different modalities, and is used to extract structured event information from this modality. For example, if the multi-modal data corresponds to text modality, visual modality, and audio modality, then the event extraction model may include a text event extraction branch corresponding to the text modality, a visual event extraction branch corresponding to the visual modality, and an audio event extraction branch corresponding to the audio modality.

[0110] It can be understood that the event extraction branches for different modalities provided in the embodiments of the present application can all adopt a multi-layer Transformer model.

[0111] Specifically, the first multi-modal data includes data of the first modality and data of the second modality. The computer device can input the data of the first modality into the event extraction branch of the first modality to obtain the first event information of the first modality, and input the data of the second modality into the event extraction branch of the second modality to obtain the first event information of the second modality. After determining the training loss corresponding to the data of the first modality based on the first event information, event labels, and cross-modal reconstruction data of the second modality corresponding to the data of the first modality, the parameters of the event extraction branch of the first modality can be adjusted based on the training loss corresponding to the data of the first modality. After determining the training loss corresponding to the data of the second modality based on the first event information, event labels, and cross-modal reconstruction data of the first modality corresponding to the data of the second modality, the parameters of the event extraction branch of the second modality can be adjusted based on the training loss corresponding to the data of the second modality. By adjusting the event extraction branches for each modality, the adjusted event extraction model is obtained, and then step S202 is returned to execute until the convergence condition is reached, and the target event extraction branches corresponding to each modality after training are obtained.

[0112] The above embodiments are described in conjunction with the following examples. For example, the first multimodal data includes text data T, visual data I, and audio data S. The event extraction model may include a text event extraction branch corresponding to the text modality, a visual event extraction branch corresponding to the visual modality, and an audio event extraction branch corresponding to the audio modality. When the computer device obtains the training loss corresponding to the text data T and the training loss corresponding to the visual data I and the training loss corresponding to the audio data S after that, based on the training loss the parameters of the text event extraction branch are optimized to obtain the target text event extraction branch. Based on the training loss the parameters of the visual event extraction branch are optimized to obtain the target visual event extraction branch. Based on the training loss the parameters of the audio event extraction branch are optimized to obtain the target audio event extraction branch.

[0113] In the above embodiments, the computer device designs independent event extraction branches for each modality. Each modality (such as text, image, audio, etc.) can be specifically processed according to its characteristics, so that the features of each modality can be more accurately extracted and modeled, enabling the model to better handle the complexity in multimodal data. By jointly optimizing different modality branches through the joint cross-modal reconstruction loss and event extraction loss, the problems of error accumulation and accuracy decline that may occur during independent training of different branches are avoided. Subsequently, when using the target event extraction branch to extract event information from multimodal data, the accuracy of event information extraction can be improved.

[0114] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the processing method of the above event extraction model further includes the following steps: fusing the first event information of each modality through the fusion network to obtain the first fused event information; determining the first fusion loss based on the first fused event information and the fused event label; optimizing the parameters of the fusion network based on the first fusion loss to obtain the target fusion network.

[0115] Among them, the fusion network is a neural network structure for fusing the event extraction results corresponding to data of multiple modalities. The fused event label is the fused event information pre-annotated for the training of the fusion network. The structure of the fused event label is similar to that of the single-modal event label, but it is the final version extracted and verified from multiple modality information and has a more complete and accurate structure.

[0116] Specifically, after obtaining the first event information of each modality, the computer device can input the first event information of each modality into the fusion network, process the first event information of each modality through the network layer of the fusion network to obtain the first fusion event information, and obtain the fusion event label corresponding to the pre-annotated multimodal data. Based on the difference between the first fusion event information and the fusion event label, the first fusion loss is determined, and the parameters of the fusion network are adjusted based on the first fusion loss to obtain the adjusted fusion network. Then, return to execute step S202 until the convergence condition is reached, and the trained target fusion network is obtained.

[0117] In the above embodiment, the computer device fuses the first event information of each modality through the fusion network, which can integrate the information from different modalities into a unified representation. The event information of different modalities can complement each other, and the fusion network can make full use of the information of each modality, thereby improving the accuracy of overall event extraction. In addition, by calculating the loss between the fusion event information and the label (such as the first fusion loss), the model can continuously adjust the fusion strategy during the optimization process, so as to better integrate the features of different modalities. Therefore, when using the target fusion network to fuse the event information extracted from different modalities to obtain the final event information later, the accuracy of the final event information can be further improved.

[0118] In one embodiment, the process of the computer device fusing the first event information of each modality through the fusion network to obtain the first fusion event information includes the following steps: obtaining the information confidence corresponding to the first event information of each modality respectively; and fusing the first event information of each modality through the fusion network based on the information confidence of the first event information of each modality to obtain the first fusion event information.

[0119] Among them, the information confidence refers to the quantitative representation of the reliability or credibility of the event information extracted in a certain modality. When the event information is represented in a sequence form, the information credibility can include the credibility of each element in the event information. For example, if the first event information Y is represented in a sequence form, Y = [E, e 1 , e 2 ,⋯, e n ] , where E is the event category, is the i-th event element, and the corresponding information confidence can also be represented by a confidence sequence of the same length as Y, such as Conf Y =[ C E , c 1 , c 2 ,⋯, c n ] , where represents the confidence level of event type E, is the confidence level of the event element . The values of each confidence level can be within the range of [0, 1].

[0120] Specifically, after the computer device obtains the first event information of each modality, it can also obtain the information confidence levels corresponding to the first event information of each modality. Then, it inputs the first event information and the information confidence levels of each modality into the fusion network. Based on the information confidence levels of the first event information of each modality, the network layer of the fusion network fuses the first event information of each modality to obtain the first fused event information. It also obtains the fused event label corresponding to the pre-annotated multi-modal data. Based on the difference between the first fused event information and the fused event label, it determines the first fusion loss, and adjusts the parameters of the fusion network based on the first fusion loss to obtain the adjusted fusion network. Then, it returns to execute step S202 until the convergence condition is reached, and obtains the trained target fusion network.

[0121] In one embodiment, the process of the computer device obtaining the information confidence levels corresponding to the first event information of each modality includes: when the event extraction branches of each modality output the first event information of the corresponding modality, they can also output the information confidence levels corresponding to the first event information; or verify the first event information of each modality to obtain the verification results of the first event information of each modality, and determine the information confidence levels corresponding to the first event information of the corresponding modality based on the verification results of the first event information of each modality; or set the corresponding information confidence levels for the first event information of each modality based on prior knowledge.

[0122] In the above embodiments, the quality and consistency of multi-modal data often vary. Some modalities may provide more explicit or powerful event features, while other modalities may contain noise or incomplete information. By considering the information confidence levels during fusion, the computer device enables the model to dynamically adjust the fusion weights of each modality according to the actual situation. This enables the model to automatically learn which modalities are more useful in specific situations, thereby improving the quality of the fused event information, that is, improving the accuracy of the final event information.

[0123] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the processing method of the above event extraction model further includes the following steps: performing event extraction processing on the second multi-modal data through the target event extraction branches of each modality to obtain second event information of each modality; fusing the second event information of each modality through the fusion network to obtain second fused event information; determining a second fusion loss based on the second fused event information and the fused event label; and optimizing the parameters of the fusion network based on the second fusion loss to obtain a target fusion network.

[0124] Among them, the second multi-modal data may specifically be multi-modal data describing the second sample event, and is also sample data for training the fusion network. The second multi-modal data is composed of multiple different modality data types, and these data types can describe the same event from different perspectives and dimensions, so as to provide more comprehensive and accurate event information. It should be noted that the second sample event in the embodiments of the present application may be the same event as the first sample event, or the second sample event may be a different event from the first sample event. Correspondingly, the second multi-modal data may be the same as the first multi-modal data, or the second multi-modal data may be different from the first multi-modal data.

[0125] Specifically, after the computer device obtains the target event extraction branches corresponding to each modality, it can input the data of each modality in the second multi-modal data into the corresponding target event extraction branches of each modality respectively to obtain second event information corresponding to the data of each modality, and input the second event information corresponding to the data of each modality into the fusion network. The network layer of the fusion network processes the second event information of each modality to obtain second fused event information, and obtains the fused event label corresponding to the pre-annotated multi-modal data. Based on the difference between the second fused event information and the fused event label, a second fusion loss is determined, and the parameters of the fusion network are adjusted based on the second fusion loss to obtain an adjusted fusion network. Then, return to execute the step of "performing event extraction processing on the second multi-modal data through the target event extraction branches of each modality to obtain second event information of each modality" until the convergence condition is reached to obtain a trained target fusion network.

[0126] In one embodiment, the process of the computer device fusing the second event information of each modality through the fusion network to obtain second fused event information includes the following steps: obtaining the information confidence levels corresponding to the second event information of each modality; and fusing the second event information of each modality through the fusion network based on the information confidence levels of the second event information of each modality to obtain second fused event information.

[0127] Specifically, after the computer device obtains the second event information of each modality, it can also obtain the information confidence corresponding to the second event information of each modality. Then, it inputs the second event information and information confidence of each modality into the fusion network. The network layer of the fusion network fuses the second event information of each modality based on the information confidence of the second event information of each modality to obtain the second fused event information. It also obtains the fused event label corresponding to the pre-annotated multimodal data, determines the second fusion loss based on the difference between the second fused event information and the fused event label, and adjusts the parameters of the fusion network based on the second fusion loss to obtain the adjusted fusion network. Then, it returns to execute the step of "performing event extraction processing on the second multimodal data for each modality through the target event extraction branch of each modality to obtain the second event information of each modality" until the convergence condition is reached, and obtains the trained target fusion network.

[0128] In one embodiment, the process of the computer device obtaining the information confidence corresponding to the second event information of each modality includes: when the target event extraction branch of each modality outputs the second event information of the corresponding modality, it can also output the information confidence corresponding to the second event information; or verifying the second event information of each modality to obtain the verification result of the second event information of each modality, and determining the information confidence corresponding to the second event information of the corresponding modality based on the verification result of the second event information of each modality; or setting the corresponding information confidence for the second event information of each modality based on prior knowledge.

[0129] In the above embodiment, the computer device trains the fusion network by introducing the second multimodal data, enabling the model to be trained and optimized on new data, improving its adaptability to unknown or new data, and being able to handle more complex data distributions and information scenarios. Additionally, by calculating the loss between the fused event information and the label (such as the first fusion loss), the model can continuously adjust the fusion strategy during the optimization process, thereby better integrating the features of different modalities. Consequently, when using the target fusion network to fuse the event information extracted from different modalities to obtain the final event information, the accuracy of the final event information can be further improved.

[0130] In one embodiment, as Figure 6 shown, the processing method of the above event extraction model further includes the process of pre-training the event extraction model, and this process specifically includes the following steps:

[0131] S602, through the initial event extraction model, perform event extraction processing on the third multimodal data for each modality to obtain the third event information of each modality.

[0132] Among them, the initial event extraction model is a model to be pre-trained for identifying and extracting event information from multi-modal data. The initial event extraction model may specifically include initial event extraction branches corresponding to each modality. The initial event extraction branches corresponding to each modality refer to independent event extraction sub-models or sub-network structures respectively designed and deployed for data of different modalities, and are used to extract structured event information from this modality.

[0133] Specifically, the third multi-modal data may be multi-modal data describing the third sample event, and is also sample data for training the fusion network. The third multi-modal data is composed of multiple different modality data types, and these data types can describe the same event from different perspectives and dimensions, so as to provide more comprehensive and accurate event information. It should be noted that the third sample event in the embodiments of the present application may be the same event as the first sample event, or the third sample event may be a different event from the first sample event. Correspondingly, the third multi-modal data may be the same as the first multi-modal data, or the third multi-modal data may be different from the first multi-modal data.

[0134] Specifically, the third multi-modal data includes data of the first modality and data of the second modality. The computer device may input the data of the first modality into the initial event extraction branch of the first modality to obtain the third event information of the first modality, and input the data of the second modality into the initial event extraction branch of the second modality to obtain the third event information of the second modality. By performing event extraction processing on the data of different modalities in the third multi-modal data, the third event information corresponding to the data of each modality is obtained.

[0135] In one embodiment, the initial event extraction branches corresponding to each modality respectively include an encoder and a decoder. For data of a certain modality, the data of this modality may be input into the event extraction branch of this modality. The encoder of the initial event extraction branch of this modality encodes the input data to obtain an encoded vector, and the encoded vector data is input into the decoder of the initial event extraction branch of this modality. The decoder decodes the input encoded vector to obtain the third event information corresponding to the data of this modality. By performing the above processing on the data of each modality, the third event information corresponding to the data of each modality can be obtained.

[0136] In one embodiment, the computer device may first perform feature processing on the third multi-modal data to obtain features of each modality, and input the features of different modalities into the initial event extraction model to be trained. Through the initial event extraction branches corresponding to each modality of the initial event extraction model, event extraction processing is respectively performed on the features of each modality to obtain the third event information of each modality.

[0137] In one embodiment, the third multimodal data includes text data, and the computer device processes the text data to obtain text features.

[0138] Specifically, the computer device can perform word segmentation on the text data to obtain a word segmentation result, remove stop words from the word segmentation result to obtain each word, and input each word into a text feature extraction model to convert each word into a numerical text feature through the text feature extraction model. The text feature extraction model can be any one of Word2Vec, GloVe, or BERT.

[0139] In one embodiment, the third multimodal data includes visual data, and the computer device processes the visual data to obtain visual features.

[0140] Specifically, the computer device can input the visual data into a visual feature extraction model to perform feature extraction on the visual data through the visual feature extraction model to obtain visual features.

[0141] Among them, the visual feature extraction model can be a model based on a convolutional neural network (CNN). The CNN automatically learns the local features and global structure of the image through multiple layers of convolution and pooling operations, thereby generating a discriminative image feature vector. After being processed by the CNN, the original pixel matrix of the image is converted into a high-level feature vector, which can better capture the key information in the image.

[0142] In one embodiment, the third multimodal data includes audio data, and the computer device processes the audio data to obtain audio features.

[0143] Among them, the audio feature can be Mel Frequency Cepstral Coefficients (MFCC). MFCC is a feature widely used in audio processing. Based on the auditory characteristics of the human ear, the audio signal is converted into frequency-related coefficients, which can reflect the spectral envelope of the audio signal and are very effective for tasks such as speech recognition and music classification.

[0144] Specifically, the computer device can perform a short-time Fourier transform (STFT) on the audio data to obtain a Fourier transform result, extract the Mel Frequency Cepstral Coefficients of the audio data based on the Fourier transform result, and use the extracted Mel Frequency Cepstral Coefficients as the audio features of the audio data.

[0145] S604, perform original-modal data reconstruction based on the third event information of each modality to obtain the original-modal reconstruction data of each modality.

[0146] Among them, data reconstruction refers to the process of generating complete data or information related thereto using a model or algorithm based on existing data or information. The data reconstruction of the original modality in the embodiments of this application refers to reconstructing the data of the modality based on the data of one modality. For example, generating text data based on text data, specifically, generating image data based on the third event information of the text modality.

[0147] Specifically, after the computer device obtains the third event information of each modality, for the third event information of the first modality, it can input it into the original modality reconstruction model corresponding to the first modality, and generate cross-modal reconstruction data of the first modality through the original modality reconstruction model corresponding to the first modality based on the third event information of the first modality; for the third event information of the second modality, it can input it into the original modality reconstruction model corresponding to the second modality, and generate original modality reconstruction data of the second modality through the original modality reconstruction model corresponding to the second modality based on the third event information of the second modality; by performing the above processing on the third event information of each modality, the original modality reconstruction data of each modality can be obtained.

[0148] Among them, the first modality can be at least one of a text modality, an audio modality, and a visual modality, and the second modality is a modality different from the first modality among the text modality, the audio modality, and the visual modality.

[0149] For example, the third multimodal data includes text data, visual data, and audio data. By performing event extraction processing on the multimodal data, the third event information corresponding to the text data, the third event information corresponding to the visual data, and the third event information corresponding to the audio data can be obtained. Based on the third event information corresponding to the text data, data reconstruction of the original text modality is performed to obtain the original modality reconstruction text data corresponding to the text data. Based on the third event information corresponding to the audio data, data reconstruction of the original audio modality is performed to obtain the original modality reconstruction audio data corresponding to the audio data. Based on the third event information corresponding to the visual data, data reconstruction of the original visual modality is performed to obtain the original modality reconstruction visual data corresponding to the visual data.

[0150] S606, optimize the parameters of the initial event extraction model based on the original modality reconstruction data of each modality and the third event information of each modality to obtain an event extraction model.

[0151] Specifically, after the computer device obtains the original modal reconstruction data of each modality and the third event information of each modality, for the data of the first modality, it can determine the pre-training loss corresponding to the data of the first modality based on the corresponding original modal reconstruction data and the third event information. For the data of the second modality, it can determine the pre-training loss corresponding to the data of the second modality based on the corresponding original modal reconstruction data and the third event information. By performing the above processing on the original modal reconstruction data and the third event information corresponding to the data of each modality, the pre-training loss corresponding to the data of each modality is obtained, and the parameters of the initial event extraction model are adjusted based on the pre-training loss corresponding to the data of each modality, so as to obtain the adjusted initial event extraction model. Then, return to execute step S602 until the convergence condition is reached, and the trained event extraction model is obtained.

[0152] Among them, the parameters of the initial event extraction model can specifically be at least one of the weights and biases in the initial event extraction model. The convergence condition refers to, during the model training process, according to a preset standard or threshold, when the training loss value or other evaluation indicators change very slightly or reach a stable state in several consecutive iterations, it is considered that the model parameters have been fully optimized and the model performance tends to be stable. At this time, the condition for terminating the training process can be any of the following: 1. The loss change amplitude is low. For example, in several consecutive epochs, the decrease amplitude of the total loss value is lower than the preset threshold; 2. The validation metrics are stable. For example, the performance metrics (such as accuracy, mean square error, etc.) on the validation set remain stable in several consecutive cycles without obvious improvement; 3. The maximum number of iterations is reached. For example, the training process runs to the preset maximum number of iterations or training rounds.

[0153] In the above embodiment, the computer device performs event extraction processing on the third multi-modal data through the initial event extraction model to obtain the third event information of each modality, performs original modal data reconstruction based on the third event information of each modality to obtain the original modal reconstruction data of each modality, and optimizes the parameters of the initial event extraction model based on the original modal reconstruction data of each modality and the third event information of each modality. The reconstructed data of the original modality provides more features for the model to learn, which can help the model better understand the role and contribution of each modality in event extraction, thereby improving the overall event extraction performance.

[0154] In one embodiment, the process by which the computer device optimizes the parameters of the initial event extraction model based on the original modality reconstruction data of each modality and the third event information of each modality to obtain the event extraction model includes the following steps: determining the original modality reconstruction loss of each modality based on the third multimodal data and the original modality reconstruction data of each modality, and determining the second event extraction loss of each modality based on the third event information of each modality and the corresponding event label; optimizing the parameters of the initial event extraction model according to the original modality reconstruction loss and the second event extraction loss to obtain the event extraction model.

[0155] Specifically, after the computer prepares to obtain the third event information of each modality, it can obtain the event label corresponding to the data of each modality. For the data of the first modality, based on the third event information and the event label corresponding to the data of the first modality, it can determine the second event extraction loss corresponding to the data of the first modality, and based on the data of the first modality and the original modality reconstruction data of the first modality, determine the original modality reconstruction loss corresponding to the data of the first modality; for the data of the second modality, based on the third event information and the event label corresponding to the data of the second modality, determine the second event extraction loss corresponding to the data of the second modality, and based on the data of the second modality and the original modality reconstruction data of the second modality, determine the original modality reconstruction loss corresponding to the data of the second modality; by performing the above processing on the third event information and the original modality reconstruction data of each modality, the second event extraction loss and the original modality reconstruction loss of each modality can be obtained; after the computer device obtains the original modality reconstruction loss corresponding to the data of each modality and the second event extraction loss corresponding to the data of each modality, for the data of the first modality, based on the corresponding original modality reconstruction loss and the second event extraction loss, it can determine the pre-training loss corresponding to the data of the first modality, and for the data of the second modality, based on the corresponding original modality reconstruction loss and the second event extraction loss, determine the pre-training loss corresponding to the data of the second modality; by performing the above processing on the original modality reconstruction loss and the second event extraction loss corresponding to the data of each modality, the pre-training loss corresponding to the data of each modality can be obtained, and the parameters of the initial event extraction model are adjusted based on the pre-training loss corresponding to the data of each modality, so as to obtain the adjusted initial event extraction model, and then return to execute step S602 until the convergence condition is reached to obtain the trained event extraction model.

[0156] The above embodiment is described in conjunction with the following example. For example, the third multimodal data includes text data T, visual data I, and audio data S, as Figure 7 shown, the text data T corresponds to the third event information , event label , original modality reconstruction text data , then based on the third event information , event label Determine the second event extraction loss corresponding to the text data T , reconstruct the text data based on the original modality and the text data T to determine the original modality reconstruction loss corresponding to the text data T ; as Figure 8 shown, the visual data I corresponds to the third event information , event label , original modality reconstructed visual data , then based on the third event information , event label determine the second event extraction loss corresponding to the visual data I , reconstruct the visual data based on the original modality and the visual data I to determine the original modality reconstruction loss corresponding to the visual data I ; as Figure 9 shown, the audio data S corresponds to the third event information , event label , original modality reconstructed audio data , then based on the third event information , event label determine the second event extraction loss corresponding to the audio data S , reconstruct the audio data based on the original modality and the audio data S to determine the original modality reconstruction loss corresponding to the audio data S .

[0157] In one embodiment, the initial event extraction model includes initial event extraction branches for each modality, and the event extraction model includes event extraction branches for each modality; the process by which the computer device performs event extraction processing on the third multi-modal data for each modality through the initial event extraction model to obtain the third event information for each modality includes the following steps: perform event extraction processing on the third multi-modal data for each modality through the initial event extraction branches for each modality to obtain the third event information for each modality; the computer device optimizes the parameters of the initial event extraction branches for the corresponding modality according to the original modality reconstruction loss and the second event extraction loss for each modality to obtain the event extraction branches for each modality.

[0158] Specifically, the third multimodal data includes the data of the first modality and the data of the second modality. The computer device can input the data of the first modality into the initial event extraction branch of the first modality to obtain the third event information of the first modality, and input the data of the second modality into the initial event extraction branch of the second modality to obtain the third event information of the second modality. After determining the pre-training loss corresponding to the data of the first modality based on the third event information corresponding to the data of the first modality, the event label, and the original modality reconstruction data of the first modality, the parameters of the initial event extraction branch of the first modality can be adjusted based on the pre-training loss corresponding to the data of the first modality. After determining the pre-training loss corresponding to the data of the second modality based on the third event information corresponding to the data of the second modality, the event label, and the original modality reconstruction data of the second modality, the parameters of the initial event extraction branch of the second modality can be adjusted based on the pre-training loss corresponding to the data of the second modality. By adjusting the initial event extraction branch of each modality, the adjusted event extraction model can be obtained, and then return to execute step S602 until the convergence condition is reached, and the event extraction branches corresponding to each trained modality can be obtained.

[0159] In the above embodiment, the computer device is optimized by combining the original modality reconstruction loss and the event extraction loss, and the model can better understand the role and contribution of each modality in event extraction, thereby improving the overall event extraction performance.

[0160] In one embodiment, after obtaining the trained target event extraction model, the computer device can also use the target event extraction model to process the event extraction task, such as Figure 10 shown, the specific process of processing the event extraction task includes the following steps:

[0161] S1002, obtain the multi-modal data to be processed.

[0162] Among them, the multi-modal data to be processed refers to the multi-modal data that describes the event to be extracted. The multi-modal data to be processed is composed of multiple different modality data types, and these data types can describe the event to be extracted from different perspectives and dimensions.

[0163] In one embodiment, when the computer device receives an event extraction request, it can parse the received event extraction request to extract the multi-modal data to be processed of the event to be extracted from the time extraction request.

[0164] In one embodiment, the computer device can also obtain the multi-modal data to be processed from the information platform. Among them, the information platform can be at least one of a social media platform, a news information platform, a city perception platform, and an enterprise-level data system.

[0165] S1004. Perform event extraction processing on the multi-modal data to be processed for each modality through the target event extraction model, and obtain the event information of each modality of the multi-modal data to be processed.

[0166] Specifically, the computer device can input the multi-modal data to be processed into the target event extraction model, and through the target event extraction branches corresponding to each modality of the target event extraction model, perform event extraction processing on the data of different modalities in the multi-modal data to be processed respectively, and obtain the event information corresponding to the data of each modality.

[0167] In one embodiment, the target event extraction branches corresponding to each modality respectively include an encoder and a decoder. For the data of a certain modality, the data of this modality can be input into the event extraction branch of this modality. The encoder of the target event extraction branch of this modality encodes the input data to obtain an encoded vector, and the encoded vector data is input into the decoder of the target event extraction branch of this modality. The decoder decodes the input encoded vector to obtain the event information corresponding to the data of this modality. By performing the above processing on the data of each modality, the event information corresponding to the data of each modality can be obtained.

[0168] In one embodiment, the computer device can first perform feature processing on the first multi-modal data to obtain the features of each modality, and input the features of different modalities into the event extraction model to be trained. Through the event extraction branches corresponding to each modality of the event extraction model, perform event extraction processing on the features of each modality respectively, and obtain the first event information of each modality.

[0169] In one embodiment, the first multi-modal data includes text data, and the computer device performs feature processing on the text data to obtain text features.

[0170] Specifically, the computer device can perform word segmentation processing on the text data to obtain the word segmentation result, remove the stop words from the word segmentation result to obtain each word, and input each word into the text feature extraction model, so as to convert each word into a numerical text feature through the text feature extraction model. The text feature extraction model can be any one of Word2Vec, GloVe or BERT.

[0171] In one embodiment, the multi-modal data to be processed includes visual data, and the computer device performs feature processing on the visual data to obtain visual features.

[0172] Specifically, the computer device can input the visual data into the visual feature extraction model, so as to perform feature extraction on the visual data through the visual feature extraction model to obtain visual features.

[0173] Among them, the visual feature extraction model can be a model based on a convolutional neural network (CNN). The CNN automatically learns the local features and global structure of an image through multiple layers of convolution and pooling operations, thereby generating a discriminative image feature vector. After being processed by the CNN, the original pixel matrix of the image is converted into a high-level feature vector, which can better capture the key information in the image.

[0174] In one embodiment, the multi-modal data to be processed includes audio data, and the computer device processes the audio data to obtain audio features.

[0175] Among them, the audio features can be Mel Frequency Cepstral Coefficients (MFCC). MFCC is a feature widely used in audio processing. Based on the auditory characteristics of the human ear, the audio signal is converted into frequency-related coefficients, which can reflect the spectral envelope of the audio signal and are very effective for tasks such as speech recognition and music classification.

[0176] Specifically, the computer device can perform a Short-Time Fourier Transform (STFT) on the audio data to obtain the Fourier transform result, extract the Mel Frequency Cepstral Coefficients of the audio data based on the Fourier transform result, and use the extracted Mel Frequency Cepstral Coefficients as the audio features of the audio data.

[0177] In one embodiment, the target event extraction model includes target event extraction branches for each modality, and the event extraction model includes event extraction branches for each modality; step S1004 specifically includes the following: through the target event extraction branches for each modality, perform event extraction processing on the multi-modal data for each modality to obtain event information for each modality.

[0178] Specifically, the multi-modal data to be processed includes data of the first modality and data of the second modality. The computer device can input the data of the first modality into the target event extraction branch of the first modality to obtain event information of the first modality, and input the data of the second modality into the target event extraction branch of the second modality to obtain event information of the second modality. By performing the above processing on the data of each modality, event information corresponding to the data of each modality can be obtained.

[0179] S1006, based on the event information of each modality of the multi-modal data to be processed, determine the event information of the multi-modal data to be processed.

[0180] Specifically, after the computer device obtains the event information of each modality, it can fuse the event information of each modality to obtain fused event information, and this fused event information is the event information of the multi-modal data to be processed.

[0181] In the above embodiments, the target event extraction model is optimized by jointly using the cross-modal reconstruction loss and the event extraction loss during the training process. The model can improve the extraction accuracy of different modalities and the consistency between modalities simultaneously, thereby enhancing the event extraction effect of multi-modal data, avoiding the problem of error accumulation and accuracy decline that may occur during independent training. Subsequently, the target event extraction model is used to perform event extraction processing on each modality of the multi-modal data to be processed, obtaining the event information of each modality of the multi-modal data to be processed. Based on the event information of each modality of the multi-modal data to be processed, the event information of the multi-modal data to be processed is determined, which can improve the accuracy of event information extraction.

[0182] In one embodiment, the target event extraction model includes a target fusion network; the process by which the computer device determines the event information of the multi-modal data to be processed based on the event information of each modality of the multi-modal data to be processed includes the following steps: obtaining the information confidence degrees corresponding to the event information of each modality; and fusing the event information of each modality through the target fusion network based on the information confidence degrees of the event information of each modality to obtain event information.

[0183] Among them, the information confidence degree refers to a quantitative representation of the reliability or credibility of the event information extracted in a certain modality. When the event information is represented in a sequence form, the information credibility can include the credibility of each element in the event information. For example, if the event information Y is represented in a sequence form, Y = [E, e 1 , e 2 ,⋯, e n ] , where E is the event category, is the i-th event element, then the corresponding information confidence degree can also be represented by a confidence degree sequence of the same length as Y, such as Conf Y =[ C E , c 1 , c 2 ,⋯, c n ] , where represents the confidence degree of the event type E, is the confidence degree of the event element , and the values of each confidence degree can be within the interval [0, 1].

[0184] Specifically, after the computer device obtains the event information of each modality, it can also obtain the information confidence corresponding to the event information of each modality, and then input the event information and information confidence of each modality into the target fusion network. Based on the information confidence of the event information of each modality, the network layer of the target fusion network fuses the event information of each modality to obtain the event information corresponding to the multi-modal data to be processed.

[0185] In one embodiment, the process by which the computer device obtains the information confidence corresponding to the event information of each modality includes: when the target event extraction branch of each modality outputs the event information of the corresponding modality, it can also output the information confidence corresponding to the event information; or verify the event information of each modality to obtain the verification result of the event information of each modality, and determine the information confidence corresponding to the event information of the corresponding modality based on the verification result of the event information of each modality; or set the corresponding information confidence for the event information of each modality based on prior knowledge.

[0186] In the above embodiment, in multi-modal data processing, there may be noise or incomplete situations in some modalities. By using information confidence, the computer device can identify which modalities have poor data quality and automatically reduce their impact on the final event information, thereby improving the robustness of the model when processing noisy and incomplete data.

[0187] In one embodiment, the process by which the computer device determines the event information of the multi-modal data to be processed based on the event information of each modality of the multi-modal data to be processed includes the following steps: obtaining the modality weight corresponding to each modality; based on the modality weights, fusing the event information of each modality to obtain the event information.

[0188] Among them, the modality weight refers to the numerical weight assigned to different modalities and used to measure the degree of their influence on the final fusion result.

[0189] It can be understood that in different business scenarios or application environments, due to the different reliability, expression ability or context dependence of information sources, the actual contributions of each modality will also change. Therefore, according to business requirements or data characteristics, different modality weights are dynamically or preset. For example, in the scenario of breaking news event extraction, the modality weights of the text modality, visual modality and audio modality can be 0.7, 0.2 and 0.1 in turn. In the scenario of disaster site monitoring, the modality weights of the text modality, visual modality and audio modality can be 0.15, 0.6 and 0.25 in turn.

[0190] Specifically, the computer device can determine the business scenario to which the multi-modal data belongs, determine the modality weight corresponding to each modality in this business scenario, and perform weighted fusion on the event information of each modality based on the modality weights to obtain the event information.

[0191] In the above embodiments, by introducing modal weights, the computer device can enable the model to dynamically adjust its contribution to the final event information according to the importance of each modality. Through the adjustment of modal weights, the fusion process can more accurately combine information from different modalities, making the final event information more in line with the actual situation, thereby improving the accuracy of the final event information.

[0192] In one embodiment, after the computer device determines the event information of the multi-modal data to be processed, the processing method of the above event extraction model further includes the following steps: adding the event information of the multi-modal data to be processed to the event information library; when receiving a retrieval request, extracting query conditions from the retrieval request; in the event information library, searching for target event information that meets the query conditions; and returning the multi-modal data corresponding to the target event information as a retrieval result to the initiator of the retrieval request.

[0193] Among them, the event information library refers to a database system or information storage system for storing event information and its associated multi-modal data. That is to say, event information and the corresponding multi-modal data are stored correspondingly in the event information library.

[0194] The retrieval request is a request for searching for multi-modal data that matches specific conditions from the event information library. The query condition is the query constraint content for filtering target event information. The modality of the query condition can specifically be at least one of the text modality, visual modality, and audio modality. That is, the user can express the retrieval intention through multiple modalities, and the computer device can perform event matching and retrieval based on any or multiple modality information. It can be understood that when the modality of the query condition is the text modality, the query condition can be text data such as a query text; when the modality of the query condition is the visual modality, the query condition can be visual data such as a query image, query video, etc.; when the modality of the query condition is the audio modality, the query condition can be audio data such as a query audio.

[0195] Specifically, after the computer device determines the event information of the multi-modal data to be processed, it can store the multi-modal data to be processed and the event information of the multi-modal data to be processed correspondingly in the event information library. When the computer device receives a retrieval request, it parses the retrieval request to obtain a parsing result, extracts query conditions from the parsing result, determines query keywords based on the query conditions, and searches for target event information that matches the query keywords in the event information library, and returns the multi-modal data corresponding to the target event information as a retrieval result to the initiator of the retrieval request.

[0196] It can be understood that when the query condition is a query text, the computer device can perform event extraction processing on the query text to obtain query keywords; when the query condition is a query image or a query video, the computer device can perform event extraction processing on the query image or the query video to obtain query keywords; when the query condition is a query audio, the computer device can perform event extraction processing on the query audio to obtain query keywords.

[0197] In the above embodiments, by centrally storing all event information in the event information library, the computer device can achieve efficient information management, facilitating subsequent data retrieval and query; through the retrieval request, the user can flexibly query and obtain relevant event information.

[0198] In one embodiment, as Figure 11 shown, a processing method for an event extraction model is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps:

[0199] S1102. Through the initial event extraction model, perform event extraction processing on each modality of the third multimodal data to obtain the third event information of each modality; based on the third event information of each modality, perform data reconstruction of the original modality to obtain the original modality reconstruction data of each modality.

[0200] S1104. Determine the original modality reconstruction loss of each modality based on the third multimodal data and the original modality reconstruction data of each modality, and determine the second event extraction loss of each modality based on the third event information of each modality and the corresponding event label; according to the original modality reconstruction loss and the second event extraction loss, optimize the parameters of the initial event extraction model to obtain the event extraction model.

[0201] S1106. Through the event extraction branches of each modality of the event extraction model, perform event extraction processing on each modality of the first multimodal data to obtain the first event information of each modality; based on the first event information of each modality, perform cross-modal data reconstruction to obtain the cross-modal reconstruction data of each modality.

[0202] S1108. Determine the cross-modal reconstruction loss of each modality based on the first multimodal data and the cross-modal reconstruction data of each modality, and determine the first event extraction loss of each modality based on the first event information of each modality and the corresponding event label; according to the cross-modal reconstruction loss and the first event extraction loss, optimize the parameters of the event extraction branch of the corresponding modality to obtain the target event extraction branch of each modality.

[0203] S1110. Through the target event extraction branches of each modality, perform event extraction processing on the second multi-modal data for each modality to obtain the second event information of each modality; fuse the second event information of each modality through the fusion network of the event extraction model to obtain the second fused event information.

[0204] S1112. Based on the second fused event information and the fused event label, determine the second fusion loss; optimize the parameters of the fusion network based on the second fusion loss to obtain the target fusion network.

[0205] S1114. Obtain the multi-modal data to be processed; perform event extraction processing on the multi-modal data to be processed for each modality through the target event extraction model to obtain the event information of each modality of the multi-modal data to be processed.

[0206] S1116. Obtain the information confidence corresponding to the event information of each modality; fuse the event information of each modality through the target fusion network based on the information confidence of the event information of each modality to obtain the event information.

[0207] The present application also provides a processing system for an event extraction model. The processing system for the event extraction model includes a multi-modal data preprocessing module, a single-modal event extraction module, a native modality reconstruction module, a cross-modal reconstruction module, a fusion module, and a parameter optimization module. Refer to Figure 12 , the processing system for the event extraction model can implement the following steps:

[0208] 1. Multi-modal data preprocessing module:

[0209] Assume that the multi-modal data used for training the model includes data in the text modality (text data), data in the visual modality (visual data), and data in the audio modality (audio data). These data jointly describe different dimensions of the same event. The text data can be represented as an ordered text sequence , the visual data can be represented as a matrix I composed of pixel points, and the audio data can be represented as a signal sequence S. For subsequent analysis and processing, these multi-modal data can be preprocessed by the native modality reconstruction verification module first. The preprocessing process can include the following steps:

[0210] For the text data, the text data can be tokenized to obtain the tokenization result, and stop words can be removed from the tokenization result to obtain each word, and each word is input into the text feature extraction model to convert each word into a numerical text feature through the text feature extraction model. The text feature extraction model can be any one of Word2Vec, GloVe, or BERT.

[0211] For visual data, the visual data can be input into a visual feature extraction model to extract features from the visual data through the visual feature extraction model, obtaining visual features.

[0212] Among them, the visual feature extraction model can be a model based on a convolutional neural network (CNN). The CNN automatically learns the local features and global structure of an image through multiple convolutional and pooling operations, thereby generating a discriminative image feature vector. After being processed by the CNN, the original pixel matrix of the image is converted into a high-level feature vector, which can better capture the key information in the image.

[0213] For audio data, short-time Fourier transform (STFT) can be performed on the audio data to obtain the Fourier transform result, and the Mel-frequency cepstral coefficients of the audio data are extracted based on the Fourier transform result, and the extracted Mel-frequency cepstral coefficients are used as the audio features of the audio data.

[0214] Through the above steps, the original data of the three modalities of text, image, and audio are successfully converted into a unified feature representation. These feature vectors not only have a lower dimension but also can better retain the key information in the original data, providing a solid foundation for subsequent multi-modal data fusion and analysis. The text feature representation, image feature representation, and video feature representation can be denoted as , , .

[0215] 2. Single-modal event extraction module:

[0216] For any one type of modal data, based on the preprocessed feature representation, use an event extraction branch (such as Transformer) to generate a preliminary extraction result of event information (such as entities, event elements, etc.).

[0217] 3. Original modal reconstruction module:

[0218] It is used to take the preliminary extraction result as the input of the original modal reconstruction model and perform reverse verification through the original modal reconstruction model. This original modal reconstruction model is also a deep learning-based model, which not only evaluates the rationality of the extraction result but also generates "pseudo-data" that matches the original modality and compares it with the original data to verify the accuracy of the extraction result.

[0219] 3. Cross-modal reconstruction module:

[0220] The goal of this stage is to utilize the complementarity between different modalities and improve the overall effect of the model through a cross-validation strategy. The specific method is to use the generation result of one modality to reconstruct the information of another modality, thereby enhancing the model's understanding and generation ability of multi-modal data, and jointly improving the accuracy and consistency of the extraction result.

[0221] 4. Fusion Module:

[0222] Take the extraction results of different modalities and their corresponding confidence levels (or probabilities) as inputs, fuse the information from different modalities to generate a more accurate and comprehensive event extraction result. Among them, these confidence levels can be obtained through the internal evaluation of the model or assigned based on external verification results or prior knowledge.

[0223] 5. Parameter Optimization Module:

[0224] 1) Determine the reconstruction loss based on the original modality reconstruction result output by the original modality reconstruction module and the data of the original modality, determine the event extraction loss based on the extraction result of the original modality and the event label of the original modality, and optimize the parameters of the initial event extraction branch based on the original modality reconstruction loss and the event extraction loss to obtain the event extraction branch.

[0225] 2) Determine the reconstruction loss based on the cross-modal reconstruction result output by the cross-modal reconstruction modality and the data of the original modality, determine the event extraction loss based on the extraction result of the original modality and the event label of the original modality, and optimize the parameters of the event extraction branch based on the cross-modal reconstruction loss and the event extraction loss to obtain the target event extraction branch;

[0226] 3) Determine the fusion loss based on the fusion result and the label fusion result output by the fusion module, and optimize the parameters of the fusion network based on the fusion loss to obtain the target fusion network. Among them, the fusion result is the result obtained by fusing the extraction results of different modalities output by the target event extraction branch.

[0227] Among them, in this process, the system will dynamically adjust the weights of different modalities: for information with a higher confidence level, the model will strengthen it to ensure its greater contribution to the final output; while for information with a lower confidence level, the model will suppress or correct it to reduce its negative impact on the overall result. This strategy ensures that the final comprehensive result can both make the most of the rich information of each modality and maintain a high level of accuracy and reliability.

[0228] Through this confidence-based multi-modal fusion and result prediction method, the system can automatically balance the weights between different modalities, reasonably utilize the complementary information of different modalities, thereby improving the overall performance of the multi-modal event extraction system. This provides effective technical support for improving the intelligent level of the system and enhancing its ability to handle complex events.

[0229] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0230] Based on the same inventive concept, an embodiment of the present application further provides an event extraction model processing device for implementing the event extraction model processing method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the event extraction model processing device provided below can refer to the limitations on the event extraction model processing method in the above text, and will not be repeated here.

[0231] In one embodiment, as Figure 13 shown, an event extraction model processing device is provided, including: an event extraction module 1302, a data reconstruction module 1304, a loss determination module 1306, and a parameter optimization module 1308, where:

[0232] The event extraction module 1302 is configured to perform event extraction processing on each modality of the first multi-modal data through the event extraction model to obtain first event information of each modality;

[0233] The data reconstruction module 1304 is configured to perform cross-modal data reconstruction based on the first event information of each modality to obtain cross-modal reconstruction data of each modality;

[0234] The loss determination module 1306 is configured to determine the cross-modal reconstruction loss of each modality based on the first multi-modal data and the cross-modal reconstruction data of each modality, and determine the first event extraction loss of each modality based on the first event information of each modality and the corresponding event label;

[0235] The parameter optimization module 1308 is configured to optimize the parameters of the event extraction model according to the cross-modal reconstruction loss and the first event extraction loss to obtain a target event extraction model.

[0236] In the above embodiments, through the event extraction model, event extraction processing is performed on the first multi-modal data for each modality to obtain the first event information for each modality. Based on the first event information for each modality, cross-modal data reconstruction is performed to obtain the cross-modal reconstruction data for each modality. Based on the first multi-modal data and the cross-modal reconstruction data for each modality, the cross-modal reconstruction loss for each modality is determined. The parameters of the event extraction model are optimized according to the cross-modal reconstruction loss, so that the model can capture the mutual relationship between different modalities, help the model understand the connection between different modalities, reduce the inconsistency between the independent extraction results of modalities, and avoid the inconsistency problem caused by independent modality training. Based on the first event information for each modality and the corresponding event labels, the first event extraction loss for each modality is determined. According to the first event extraction loss, the parameters of the event extraction model are optimized, which can improve the accuracy of the model for event extraction. Therefore, by jointly optimizing the cross-modal reconstruction loss and the event extraction loss, the model can simultaneously improve the extraction accuracy of different modalities and the consistency between modalities, thereby enhancing the event extraction effect of multi-modal data, avoiding the problems of error accumulation and accuracy decline that may occur during independent training. Furthermore, when using the target event extraction model to extract event information from multi-modal data subsequently, the accuracy of event information extraction can be improved.

[0237] In one embodiment, the event extraction model includes event extraction branches for each modality, and the target event extraction model includes target event extraction branches for each modality. The event extraction module 1302 is further configured to: through the event extraction branches for each modality, perform event extraction processing on the first multi-modal data to obtain the first event information for each modality. The parameter optimization module 1308 is further configured to: according to the cross-modal reconstruction loss and the first event extraction loss for each modality, optimize the parameters of the event extraction branches for the corresponding modality to obtain the target event extraction branches for each modality.

[0238] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; as Figure 14 shown, the apparatus further includes: a fusion module 1310, configured to fuse the first event information for each modality through the fusion network to obtain the first fused event information; a loss determination module 1306, further configured to determine the first fusion loss based on the first fused event information and the fused event labels; the parameter optimization module 1308 is further configured to optimize the parameters of the fusion network based on the first fusion loss to obtain the target fusion network.

[0239] In one embodiment, the fusion module 1310 is further configured to: obtain the information confidence corresponding to the first event information of each modality; and fuse the first event information of each modality based on the information confidence of the first event information of each modality through a fusion network to obtain first fused event information.

[0240] In one embodiment, the event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the event extraction module 1302 is further configured to perform event extraction processing on the second multi-modal data through the target event extraction branches of each modality to obtain the second event information of each modality; the apparatus further includes a fusion module 1310 configured to fuse the second event information of each modality through the fusion network to obtain second fused event information; the loss determination module 1306 is further configured to determine a second fusion loss based on the second fused event information and the fused event label; and the parameter optimization module 1308 is further configured to optimize the parameters of the fusion network based on the second fusion loss to obtain the target fusion network.

[0241] In one embodiment, as Figure 14 shown, the apparatus further includes a pre-training module 1312 configured to: perform event extraction processing on the third multi-modal data through an initial event extraction model to obtain the third event information of each modality; perform original modality data reconstruction based on the third event information of each modality to obtain the original modality reconstruction data of each modality; and optimize the parameters of the initial event extraction model based on the original modality reconstruction data of each modality and the third event information of each modality to obtain the event extraction model.

[0242] In one embodiment, the pre-training module 1312 is further configured to: determine the original modality reconstruction loss of each modality based on the third multi-modal data and the original modality reconstruction data of each modality, and determine the second event extraction loss of each modality based on the third event information of each modality and the corresponding event label; and optimize the parameters of the initial event extraction model according to the original modality reconstruction loss and the second event extraction loss to obtain the event extraction model.

[0243] In one embodiment, as Figure 14 shown, the apparatus includes an inference module 1314 configured to: obtain multi-modal data to be processed; perform event extraction processing on the multi-modal data to be processed through the target event extraction model to obtain the event information of each modality of the multi-modal data to be processed; and determine the event information of the multi-modal data to be processed based on the event information of each modality of the multi-modal data to be processed.

[0244] In one embodiment, the target event extraction model includes a target fusion network and target event extraction branches for each modality; the inference module 1314 is further configured to: obtain the information confidence corresponding to the event information of each modality; and fuse the event information of each modality based on the information confidence of the event information of each modality through the target fusion network to obtain event information.

[0245] In one embodiment, the inference module 1314 is further configured to: obtain the modality weights corresponding to each modality; and fuse the event information of each modality based on the modality weights to obtain event information.

[0246] In one embodiment, the apparatus further includes an event retrieval module 1316, configured to: add the event information of the multi-modal data to be processed to the event information library; extract a query condition from the retrieval request when receiving the retrieval request; search for target event information that meets the query condition in the event information library; and return the multi-modal data corresponding to the target event information to the initiator of the retrieval request as a retrieval result.

[0247] Each module in the processing apparatus of the above event extraction model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0248] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 15As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a processing method of an event extraction model. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0249] Those skilled in the art can understand that Figure 15 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0250] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0251] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0252] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0253] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0254] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, etc., and are not limited thereto.

[0255] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0256] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A processing method for an event extraction model, characterized in that The method includes: Performing event extraction processing on the first multimodal data for each modality through an event extraction model to obtain first event information for each of the modalities; Performing cross-modal data reconstruction based on the first event information for each of the modalities to obtain cross-modal reconstruction data for each of the modalities; Determining cross-modal reconstruction losses for each of the modalities based on the first multimodal data and the cross-modal reconstruction data for each of the modalities, and determining first event extraction losses for each of the modalities based on the first event information for each of the modalities and corresponding event labels; Optimizing the parameters of the event extraction model according to the cross-modal reconstruction losses and the first event extraction losses to obtain a target event extraction model.

2. The method according to claim 1, wherein The event extraction model includes event extraction branches for each of the modalities, and the target event extraction model includes target event extraction branches for each of the modalities; The performing, through the event extraction model, event extraction processing on the first multimodal data for each modality to obtain first event information for each of the modalities includes: Performing event extraction processing on the first multimodal data for each modality through the event extraction branches for each of the modalities to obtain first event information for each of the modalities; The optimizing the parameters of the event extraction model according to the cross-modal reconstruction losses and the first event extraction losses to obtain a target event extraction model includes: Optimizing the parameters of the event extraction branches for the corresponding modalities according to the cross-modal reconstruction losses and the first event extraction losses for each of the modalities to obtain target event extraction branches for each of the modalities.

3. The method according to claim 2, wherein The event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the method further includes: Fusing the first event information for each of the modalities through the fusion network to obtain first fused event information; Determining a first fusion loss based on the first fused event information and a fused event label; Optimizing the parameters of the fusion network based on the first fusion loss to obtain a target fusion network.

4. The method according to claim 3, characterized in that The fusing, through the fusion network, the first event information for each of the modalities to obtain first fused event information includes: Obtaining information confidence levels respectively corresponding to the first event information for each of the modalities; Fusing the first event information for each of the modalities through the fusion network based on the information confidence levels of the first event information for each of the modalities to obtain first fused event information.

5. The method according to claim 2, wherein The event extraction model further includes a fusion network; the target event extraction model further includes a target fusion network; the method further includes: Performing event extraction processing on the second multimodal data for each modality through the target event extraction branches for each of the modalities to obtain second event information for each of the modalities; Fusing the second event information for each of the modalities through the fusion network to obtain second fused event information; Determining a second fusion loss based on the second fused event information and a fused event label; Optimizing the parameters of the fusion network based on the second fusion loss to obtain a target fusion network.

6. The method according to claim 1, wherein The method further includes: Through the initial event extraction model, perform event extraction processing on the third multimodal data for each of the modalities to obtain third event information for each of the modalities; Based on the third event information for each of the modalities, perform data reconstruction of the original modality to obtain original modality reconstruction data for each of the modalities; Based on the original modality reconstruction data for each of the modalities and the third event information for each of the modalities, optimize the parameters of the initial event extraction model to obtain the event extraction model.

7. The method according to claim 6, characterized in that, The optimizing the parameters of the initial event extraction model based on the original modality reconstruction data for each of the modalities and the third event information for each of the modalities to obtain the event extraction model includes: Determine the original modality reconstruction loss for each of the modalities based on the third multimodal data and the original modality reconstruction data for each of the modalities, and determine the second event extraction loss for each of the modalities based on the third event information for each of the modalities and the corresponding event labels; According to the original modality reconstruction loss and the second event extraction loss, optimize the parameters of the initial event extraction model to obtain the event extraction model.

8. The method according to any one of claims 1 to 7, characterized in that, The method includes: Obtain the multimodal data to be processed; Through the target event extraction model, perform event extraction processing on the multimodal data to be processed for each modality to obtain event information for each modality of the multimodal data to be processed; Based on the event information for each modality of the multimodal data to be processed, determine the event information of the multimodal data to be processed.

9. The method according to claim 8, wherein The target event extraction model includes a target fusion network; the determining the event information of the multimodal data to be processed based on the event information for each modality of the multimodal data to be processed includes: Obtain the information confidence corresponding to the event information for each of the modalities; Through the target fusion network, based on the information confidence of the event information for each modality, fuse the event information for each modality to obtain event information.

10. The method according to claim 8, wherein The determining the event information of the multimodal data to be processed based on the event information for each modality of the multimodal data to be processed includes: Obtain the modality weights corresponding to each of the modalities; Based on the modality weights, fuse the event information for each modality to obtain event information.

11. The method according to claim 8, wherein After determining the event information of the multimodal data to be processed, the method further includes: Add the event information of the multimodal data to be processed to the event information library; When a retrieval request is received, extract the query conditions from the retrieval request; In the event information library, search for target event information that meets the query conditions; Return the multimodal data corresponding to the target event information as the retrieval result to the initiator of the retrieval request.

12. A processing device for an event extraction model, characterized in that, The apparatus includes: An event extraction module, configured to perform event extraction processing on the first multimodal data for each modality through an event extraction model to obtain first event information for each modality; A data reconstruction module, configured to perform cross-modal data reconstruction based on the first event information for each modality to obtain cross-modal reconstruction data for each modality; A loss determination module, configured to determine the cross-modal reconstruction loss of each modality based on the first multi-modal data and the cross-modal reconstruction data of each modality, and determine the first event extraction loss of each modality based on the first event information of each modality and the corresponding event label; A parameter optimization module, configured to optimize the parameters of the event extraction model according to the cross-modal reconstruction loss and the first event extraction loss, so as to obtain a target event extraction model.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product, comprising a computer program, characterized in that, When this computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.