An event analysis method, system and device based on generation tasks and multi-modalities

By adopting an event analysis method based on generative tasks and multimodal in urban event analysis, and using a generative large model and a cross-modal attention mechanism for feature fusion and optimization, the problem of difficult to accurately analyze event directions and emotional dynamics in the existing technology is solved, and more accurate and timely urban event analysis and governance decisions are achieved.

CN119557603BActive Publication Date: 2025-05-30CETC BIGDATA RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510116350.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing feature fusion technology based on deep learning is difficult to accurately grasp the event trend and emotional dynamics in urban event analysis, which affects the scientificity and timeliness of urban governance decisions.

Method used

An event analysis method based on generation tasks and multimodality is adopted. By obtaining multimodal data (text, images, audio), using the generation model to generate structured event information, construct event maps, and feature fusion and optimization are performed through cross-modal attention mechanism and graph neural network to generate emotional feature vectors and event prediction maps.

Benefits of technology

It realizes more accurate urban event analysis, can capture event characteristics and emotional dynamics more comprehensively, improve the scientificity and timeliness of urban governance decisions, and avoid missing the best emergency response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557603B_ABST
    Figure CN119557603B_ABST
Patent Text Reader

Abstract

The present application discloses an event analysis method, system and device based on generation tasks and multi-modalities, which are used to quickly and accurately grasp the development trend of events. The method of the present application includes: generating structured event information from multi-modal data and obtaining event relationships; constructing an event graph; extracting multi-modal features and using a cross-modal attention mechanism to obtain a multi-modal sentiment relationship vector; based on a preset graph neural network and combining with the event graph to obtain a sentiment feature vector; calculating a contrastive learning loss value and obtaining a target sentiment feature based on the contrastive learning loss value; using the cross-modal attention mechanism to generate a target event graph; extracting relationship features and combining attention weights and a preset time series model to obtain an event prediction graph; extracting causal chains and reasoning, and obtaining trigger conditions and emotional factors after reasoning; combining the event prediction graph with the target event graph, node feature sequences, trigger conditions and emotional factors to obtain an event analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data analysis, and in particular to an event analysis method, system and device based on generation tasks and multi-modalities. Background Art

[0002] With the rise of the wave of smart city construction, the scale and complexity of cities have been continuously expanding, and the governance demands for aspects such as security, services, and refined management have become increasingly strong. Urban event situation monitoring and analysis have become key supporting elements, and each department hopes to combine new technologies to govern urban operations and respond to governance needs in a timely manner. However, driven by mobile Internet and Internet of Things technologies, urban event data presents complex characteristics such as polymorphism, heterogeneity, spatio-temporality, and sociality, and the formats of urban event data are diverse, bringing huge challenges to data fusion analysis.

[0003] In the prior art, event analysis is usually achieved by using deep learning-based feature fusion technology. First, the BERT model is used to extract the semantic features of text data, convert the text data into vector representations, capture the context information and semantic relationships in the text, and then a convolutional neural network is used to extract the visual features of images, and an audio processing model is used to extract the acoustic features of audio. After extraction, the extracted different-modal features are fused, and the fused features are input into a regression model for analysis and prediction to obtain event analysis results.

[0004] However, for the event analysis achieved by deep learning-based feature fusion technology, since only the basic features of each modality are fused, it is difficult to accurately grasp the event trend and emotional dynamics in the face of the complex and changeable urban events, thereby affecting the scientificity and timeliness of urban governance decisions, resulting in missing the best emergency event handling opportunity and expanding the scope of event influence. Summary of the Invention

[0005] To solve the above technical problems, this application provides an event analysis method, system and device based on generation tasks and multi-modalities.

[0006] The technical solutions provided in this application are described below:

[0007] In the first aspect of this application, an event analysis method based on generation tasks and multi-modalities is provided, and the method includes:

[0008] Obtain multi-modal data, where the multi-modal data includes text data, image data, and audio data;

[0009] Generate a task based on a preset large language model setting, input the multi-modal data into the generated task to generate structured event information, and perform relation extraction and semantic analysis on the structured event information to obtain event relations;

[0010] Construct an event graph, use the structured event information as the nodes of the event graph, and connect the nodes according to the event relations;

[0011] Extract the text features, image features, and audio features of the multi-modal data, and use a cross-modal attention mechanism to fuse the text features, image features, and audio features to obtain a multi-modal sentiment relation vector;

[0012] Aggregate the multi-modal sentiment relation vector based on a preset graph neural network, and combine it with the event graph to obtain a sentiment feature vector;

[0013] Construct positive sample pairs and negative sample pairs according to the sentiment feature vector and the multi-modal data. Based on the positive sample pairs and the negative sample pairs, use a preset contrast function to calculate the contrast learning loss value, and optimize the alignment of the sentiment feature vector and the multi-modal data through the contrast learning loss value to obtain a target sentiment feature;

[0014] Transmit the sentiment feature vector to the nodes of the event graph, and use the cross-modal attention mechanism to adjust the attention weights of the sentiment feature vector to generate a target event graph;

[0015] Use the preset graph neural network to extract the relation features in the target event graph, combine them with the attention weights to obtain a node feature sequence, and input the node feature sequence into a preset time series model to capture the dynamic evolution law, and obtain an event prediction graph according to the dynamic evolution law;

[0016] Extract the causal chain in the event prediction graph, and perform causal relationship reasoning based on the node feature sequence and the causal chain. After reasoning, obtain the triggering conditions and emotional factors;

[0017] Match the event prediction graph with the target event graph, and calculate the event analysis result by combining the node feature sequence, the triggering conditions, and the emotional factors.

[0018] Optionally, the extracting the text features, image features, and audio features of the multi-modal data, and using a cross-modal attention mechanism to fuse the text features, image features, and audio features to obtain a multi-modal sentiment relation vector includes:

[0019] Extract the text semantics and text sentiment in the multimodal data using a pre-set language large model, and obtain text features based on the text semantics and the text sentiment;

[0020] Extract the image scene and image expression in the multimodal data using ResNet or Vision Transformer, and obtain image features based on the image scene and the image expression;

[0021] Extract the audio intonation, audio volume, and audio rhythm in the multimodal data using the Wav2Vec audio model, and obtain audio features based on the audio intonation, the audio volume, and the audio rhythm;

[0022] Use a cross-modal attention mechanism to calculate the target weights of the text features, the image features, and the audio features, and combine the target weights with the text features, the image features, and the audio features to obtain a multimodal sentiment relationship vector.

[0023] Optionally, set a generation task based on a pre-set generation large model, and input the multimodal data into the generation task. The generated structured event information includes:

[0024] Set a generation task based on a pre-set generation large model, input the multimodal data into the generation task for training, and obtain a structured description;

[0025] Under the pre-set parameters of the current pre-set generation large model, combine the multimodal data and the corresponding structured description into a sample pair, and calculate the target probability of the structured event information generated by the input multimodal data based on the sample pair;

[0026] According to the target probability, adjust the pre-set parameters in combination with the cross-entropy loss function, and generate structured event information after adjustment.

[0027] Optionally, the adjustment of the pre-set parameters according to the target probability in combination with the cross-entropy loss function and the generation of structured event information after adjustment are represented by the following formula:

[0028] ;

[0029] Among them, represents the structured event information, represents the quantity of the multimodal data, represents the target probability, represents the pre-set parameters, represents the th input multimodal data; represents the The structured description of the output.

[0030] Optionally, after the event prediction graph is matched with the target event graph, and the event analysis result is calculated by combining the node feature sequence, the trigger condition, and the emotion factor, the method further includes:

[0031] Converting the data structure of the event analysis result into a target data structure, and generating a visual analysis result of the converted event analysis result by using a preset visualization tool.

[0032] Optionally, the causal chain in the event prediction graph is extracted, and causal relationship reasoning is performed based on the node feature sequence and the causal chain. The trigger condition and the emotion factor obtained after the reasoning are represented by the following formula:

[0033] ;

[0034] Wherein, represents the trigger condition, represents the emotion factor, represents the corresponding trigger condition at the time step, represents the node feature sequence, represents the weight matrix parameter, represents the causal chain.

[0035] Optionally, the contrast learning loss value is represented by the following formula:

[0036] ;

[0037] Wherein, represents the contrast loss value, represents the emotional feature vector of the positive sample pair, represents the emotional feature vector of the negative sample pair, represents an adjustable parameter, represents the number of the multimodal data, represents the similarity between the emotional feature vectors of the positive sample pair and the negative sample pair, represents the traversal index variable.

[0038] The second aspect of the present application provides an event analysis system based on a generation task and multimodality. The system includes:

[0039] An acquisition unit, configured to acquire multimodal data, where the multimodal data includes text data, image data, and audio data;

[0040] The first generation unit is used to generate a task based on a preset generation large model setting, input the multimodal data into the generation task, generate structured event information, and perform relation extraction and semantic analysis on the structured event information to obtain event relations;

[0041] The construction unit is used to construct an event graph, use the structured event information as nodes of the event graph, and connect the nodes according to the event relations;

[0042] The fusion unit is used to extract the text features, image features, and audio features of the multimodal data, and use a cross-modal attention mechanism to fuse the text features, the image features, and the audio features to obtain a multimodal sentiment relation vector;

[0043] The aggregation unit is used to aggregate the multimodal sentiment relation vector based on a preset graph neural network, and combine it with the event graph to obtain a sentiment feature vector;

[0044] The first calculation unit is used to construct positive sample pairs and negative sample pairs according to the sentiment feature vector and the multimodal data, based on the positive sample pairs and the negative sample pairs, use a preset contrast function to calculate a contrast learning loss value, and optimize the alignment of the sentiment feature vector and the multimodal data through the contrast learning loss value to obtain a target sentiment feature;

[0045] The second generation unit is used to transmit the sentiment feature vector to the nodes of the event graph, and use the cross-modal attention mechanism to adjust the attention weights of the sentiment feature vector to generate a target event graph;

[0046] The capture unit is used to extract relation features in the target event graph using the preset graph neural network, combine the attention weights to obtain a node feature sequence, and input the node feature sequence into a preset time series model to capture dynamic evolution rules, and obtain an event prediction graph according to the dynamic evolution rules;

[0047] The inference unit is used to extract the causal chain in the event prediction graph, and perform causal relation inference based on the node feature sequence and the causal chain, and obtain a trigger condition and an emotion factor after inference;

[0048] The second calculation unit is used to match the event prediction graph with the target event graph, and combine the node feature sequence, the trigger condition, and the emotion factor to calculate an event analysis result.

[0049] Optionally, the fusion unit is specifically used for:

[0050] Extract the text semantics and text sentiment in the multimodal data using a pre-set language large model, and obtain text features based on the text semantics and the text sentiment;

[0051] Extract the image scene and image expression in the multimodal data using ResNet or Vision Transformer, and obtain image features based on the image scene and the image expression;

[0052] Extract the audio intonation, audio volume, and audio rhythm in the multimodal data using the Wav2Vec audio model, and obtain audio features based on the audio intonation, the audio volume, and the audio rhythm;

[0053] Use a cross-modal attention mechanism to calculate the target weights of the text features, the image features, and the audio features, and combine the target weights with the text features, the image features, and the audio features to obtain a multimodal sentiment relationship vector.

[0054] Optionally, the first generation unit is specifically configured to:

[0055] Set a generation task based on a pre-set generation large model, input the multimodal data into the generation task for training, and obtain a structured description;

[0056] Under the pre-set parameters of the current pre-set generation large model, combine the multimodal data and the corresponding structured description into a sample pair, and calculate the target probability of the structured event information generated by the input multimodal data based on the sample pair;

[0057] According to the target probability, adjust the pre-set parameters in combination with the cross-entropy loss function, and generate structured event information after adjustment.

[0058] Optionally, the adjusting the pre-set parameters in combination with the cross-entropy loss function according to the target probability and generating structured event information after adjustment is represented by the following formula:

[0059] ;

[0060] Where represents the structured event information, represents the quantity of the multimodal data, represents the target probability, represents the pre-set parameters, represents the th input multimodal data; represents the th output structured description.

[0061] Optionally, it further includes a third generation unit, specifically used for:

[0062] Converting the data structure of the event analysis result into a target data structure, and using a preset visualization tool to generate a visualization analysis result from the converted event analysis result.

[0063] Optionally, extracting the causal chain in the event prediction graph, performing causal relationship reasoning based on the node feature sequence and the causal chain, and obtaining a trigger condition and an emotional factor after reasoning, which are represented by the following formula:

[0064] ;

[0065] Wherein, represents the trigger condition, represents the emotional factor, represents the corresponding trigger condition in the time step, represents the node feature sequence, represents the weight matrix parameter, represents the causal chain.

[0066] Optionally, the contrast learning loss value is represented by the following formula:

[0067] ;

[0068] Wherein, represents the contrast loss value, represents the emotional feature vector of the positive sample pair, represents the emotional feature vector of the negative sample pair, represents an adjustable parameter, represents the number of the multimodal data, represents the similarity between the emotional feature vectors of the positive sample pair and the negative sample pair, represents the traversal index variable.

[0069] The third aspect of the present application provides an event analysis device based on a generation task and multimodality, and the device includes:

[0070] A processor, a memory, an input / output unit, and a bus;

[0071] The processor is connected to the memory, the input / output unit, and the bus;

[0072] The memory stores a program, and the processor calls the program to execute the method in the first aspect and any optional method in the first aspect.

[0073] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the method of the first aspect and any optional method in the first aspect.

[0074] As can be seen from the above technical solutions, the present application has the following advantages:

[0075] Based on the generation task and multi-modal event analysis method, by using multi-modal data, large model generation tasks, and cross-modal attention mechanisms, it solves the limitation that the existing deep learning-based feature fusion technology only fuses the basic features of each modality, and realizes more accurate urban event analysis. First, extract the text, image, audio, and video features of multi-modal data, and use the cross-modal attention mechanism to dynamically allocate weights to form a multi-modal sentiment relationship vector, which can more comprehensively and deeply explore the potential associations between data and improve the ability to capture event features. Then, by generating structured event information and constructing an event graph, combined with a preset graph neural network to aggregate the multi-modal sentiment relationship vector, generate a sentiment feature vector, and closely combine the sentiment information of the event with the structure and relationship to more accurately grasp the event trend and sentiment dynamics.

[0076] After generating the sentiment feature vector, use the contrastive learning loss value calculated by the preset contrastive loss function to optimize the sentiment feature vector and multi-modal data. Then, input the sentiment feature vector as the node attribute of the event graph into the nodes, and use the cross-modal attention mechanism to track the development and changes of the event in real time, dynamically adjust the weight of the sentiment feature vector, and then generate the target event graph. Subsequently, use the preset graph neural network to process the target event graph and extract the relationship features of the nodes therein. During the extraction process, the graph neural network can deeply explore the potential associations between nodes and provide reliable relationship information for subsequent analysis. After extracting the relationship features, combine the attention weights to obtain the node feature sequence, and input the node feature sequence into the preset time series model. After receiving the node feature sequence, the time series model can capture the dynamic evolution law of the event in time series through iterative calculation, generate an event prediction graph based on the dynamic evolution law, and thus predict the future development trend of the event. Finally, fuse the event prediction graph with the target event graph and combine the node feature sequence to obtain the event analysis result. The event analysis result obtained by the present application comprehensively considers the current state and future trend of the event, and can provide a more comprehensive and accurate decision-making basis in the face of the complex and changeable urban events, which helps to improve the scientificity and timeliness of urban governance decisions, avoid missing the best emergency event handling opportunity, and thus effectively control the influence range of the event. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] To more clearly illustrate the technical solutions in this application, the following will briefly introduce the attached drawings required for the description of the embodiments. Obviously, the attached drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other attached drawings can be obtained based on these drawings.

[0078] Figure 1 Schematic flowchart of an embodiment of the event analysis method based on generation tasks and multi-modal for this application;

[0079] Figure 2 Schematic flowchart of another embodiment of the event analysis method based on generation tasks and multi-modal for this application;

[0080] Figure 3 Schematic structural diagram of an embodiment of the event analysis system based on generation tasks and multi-modal for this application;

[0081] Figure 4 Schematic structural diagram of an embodiment of the event analysis device based on generation tasks and multi-modal for this application. Detailed implementation manners

[0082] This application provides an event analysis method based on generation tasks and multi-modal, which can quickly and accurately grasp the development trend of events. It should be noted that the event analysis method based on generation tasks and multi-modal of this application is applied to a terminal.

[0083] Please refer to Figure 1 , this application first provides an embodiment of the event analysis method based on generation tasks and multi-modal, and this embodiment includes:

[0084] S101. Obtain multi-modal data, where the multi-modal data includes text data, image data, and audio data;

[0085] In this embodiment, in order to comprehensively obtain the multi-modal data of the urban event situation, it is necessary to separately screen the source platforms of the text data, image data, and audio data. The comprehensive acquisition of the multi-modal data can provide a richer information source for subsequent analysis.

[0086] Text data is mainly obtained by web crawling technology from various news websites, social media platforms, and government bulletin boards to capture text information related to urban events. The text information covers event descriptions, public discussions, etc.; image data is mainly obtained by real-time intercepting key frames from surveillance cameras and capturing pictures related to urban events on social platforms. The intercepted key frames and relevant pictures on social platforms can reflect the on-site situation when urban events occur; audio data is mainly obtained by setting up recording devices in public places, audio from on-site interviews, etc., which can capture on-site sound information, such as people's shouts, alarm sounds, etc., to assist in event judgment from multiple dimensions. Through multiple channels and methods, it is ensured that rich and comprehensive multimodal data is obtained, improving the accuracy and comprehensiveness of urban situation monitoring and avoiding the limitations of single-modal data.

[0087] S102. Set a generation task based on a preset generation large model, input the multimodal data into the generation task to generate structured event information, and perform relation extraction and semantic analysis on the structured event information to obtain event relations;

[0088] When setting a generation task based on a preset generation large model, preprocess the obtained multimodal data. The preprocessing includes removing noise, filtering out irrelevant information, and handling missing values, and then input the preprocessed multimodal data into the generation task. The preset generation large model in the generation task generates structured event information.

[0089] After generating the structured event information, the preset generation large model will also perform relation extraction and semantic analysis on the generated structured event information to identify causal relationships, temporal sequence relationships, etc. between events. After the identification is completed, event relations are obtained. Performing relation extraction and semantic analysis on the structured event information further explores the internal connections between events, making the understanding of the urban situation more in-depth and comprehensive.

[0090] S103. Construct an event graph, use the structured event information as the nodes of the event graph, and connect the nodes according to the event relations;

[0091] When constructing the event graph, use the processed structured event information as the nodes of the event graph. Each node represents a specific event, including detailed information about the event, such as event type, key elements, etc. Then, according to the event relations obtained from relation extraction and semantic analysis, establish relation connections between each node. The constructed event graph can quickly understand the relationships between events and the overall situation.

[0092] In practical applications, the event graph can be formalized as a directed graph , where is the set of event nodes, is the causal relationship between events, such as nodes and , yes Because of this, there is an edge , indicating that from point to causal relationship.

[0093] S104, extracting text features, image features, and audio features of the multimodal data, and fusing the text features, image features, and audio features using a cross-modal attention mechanism to obtain a multimodal sentiment relationship vector;

[0094] In this embodiment, text features, image features, and audio features are extracted for the acquired multimodal data, respectively, and features of different modal data are extracted respectively, so as to fully mine the key information in each modal data.

[0095] After feature extraction, the cross-modal attention mechanism is used for feature fusion. The cross-modal attention mechanism allows features of different modalities to refer to each other. For example, text features will refer to the information carried by image features and audio features during the fusion process. When analyzing urban safety events, the text feature is "the atmosphere on the scene is tense", the image feature shows a chaotic scene, and the audio hears noisy sounds. The cross-modal attention mechanism will combine the information of these different modalities and finally obtain a multimodal sentiment relationship vector that integrates information from multiple modalities. The use of the cross-modal attention mechanism breaks the information barrier between different modalities, allowing features of different modalities to complement and fuse with each other. The fused multimodal sentiment relationship vector contains rich information from multiple data sources, greatly improving the understanding and representation of the sentiment relationship of multimodal data. Compared with the feature analysis of a single modality, it can more comprehensively and accurately capture the sentiment and semantic information behind the data, providing a solid data foundation for subsequent sentiment analysis and event understanding.

[0096] S105, aggregating the multimodal sentiment relationship vector based on a preset graph neural network, and obtaining a sentiment feature vector in combination with the event graph;

[0097] The preset graph neural network takes the multimodal sentiment relationship vector as input, and uses the network structure and algorithm in the preset graph neural network to model the sentiment relationship between different modalities represented by the multimodal sentiment relationship vector. In the modeling process, the preset graph neural network regards the multimodal sentiment relationship vector as nodes and edges in the graph, and aggregates the multimodal sentiment relationship vector through graph convolution. The preset graph neural network can effectively aggregate the multimodal sentiment relationship vector and mine the hidden relationships and patterns between different modal emotions.

[0098] Subsequently, the aggregation result is combined with the already constructed event graph, which contains rich event information and the relationships between events. By integrating the aggregated multi-modal sentiment relationship vector into the event graph, the event graph not only contains information about the events themselves but also incorporates the sentiment information reflected by the multi-modal data, thereby obtaining a sentiment feature vector.

[0099] In the actual analysis of urban security events, the event graph records information such as the occurrence time, location, and relevant personnel of the events, while the multi-modal sentiment relationship vector reflects the public's sentiment towards the events. After combining the two, the sentiment feature vector can more comprehensively reflect the events and the sentiment states related to the events, understand the various situations of events in the city more deeply, help discover potential social problems and risks in advance, and provide a more comprehensive and valuable reference basis for urban governance and decision-making.

[0100] S106. Construct positive sample pairs and negative sample pairs based on the sentiment feature vector and multi-modal data. Based on the positive sample pairs and negative sample pairs, use a preset contrast function to calculate the contrast learning loss value, and optimize the alignment of the sentiment feature vector and multi-modal data through the contrast learning loss value to obtain the target sentiment feature;

[0101] Constructing positive sample pairs and negative sample pairs can learn the differences between different sentiment categories and the similarities of multi-modal data under the same sentiment category. First, combine the multi-modal data of the same sentiment category to form positive sample pairs. For example, in urban security events, all text, image, and audio data expressing the sentiment of "worry" are paired with each other to form positive sample pairs. Then, combine the multi-modal data of different sentiment categories to form negative sample pairs. For example, the multi-modal data expressing the sentiment of "anger" is combined with the multi-modal data expressing the sentiment of "calm" to form negative sample pairs.

[0102] After constructing the positive sample pairs and negative sample pairs, use a preset contrast function to calculate the contrast learning loss value, which can quantify the degree of difference between different sample pairs. Then, optimize the alignment of the sentiment feature vector and multi-modal data according to the contrast learning loss value. During the optimization process, continuously adjust the parameters to make the multi-modal features of the same sentiment category closer in the feature space and the multi-modal features of different sentiment categories farther away in the feature space, and finally obtain the target sentiment feature. The obtained target sentiment feature accurately reflects the actual sentiment in the multi-modal data, can improve the accuracy of sentiment analysis, and further provide more reliable results for the analysis of the public's sentiment towards urban security events.

[0103] Specifically, the contrast learning loss value is represented by the following formula:

[0104] ;

[0105] where, represents the contrast loss value, represents the sentiment feature vector of the positive sample pair, represents the sentiment feature vector of the negative sample pair, represents an adjustable parameter, represents the quantity of multimodal data, represents the similarity between the sentiment feature vectors of the positive sample pair and the negative sample pair, represents the traversal index variable.

[0106] S107. Transmit the sentiment feature vector into the nodes of the event graph, and use the cross-modal attention mechanism to adjust the attention weights of the sentiment feature vector to generate the target event graph;

[0107] In this embodiment, the obtained sentiment feature vector is transmitted into each node of the event graph. The nodes of the event graph respectively represent different event elements, such as the subject, time, location, behavior, etc. The sentiment feature vector carries the sentiment information contained in the multimodal data. After integrating the sentiment feature vector into the nodes, the information dimension of the event graph becomes richer, expanding from a single event structure description to a comprehensive representation including sentiment information.

[0108] Subsequently, the cross-modal attention mechanism will dynamically allocate attention weights according to the importance of different modal data for event understanding and sentiment expression. For example, when the information conveyed by the image is more important for sentiment judgment, the cross-modal attention mechanism will correspondingly increase the weight of the sentiment feature vector corresponding to the image feature; when the text is more important, it will correspondingly increase the weight of the sentiment feature vector corresponding to the text feature. Through this dynamic adjustment, it is possible to flexibly highlight the sentiment information of the key modality according to the characteristics of different events, so that each node in the event graph can fuse the multimodal sentiment information with the most appropriate weight, and finally generate the target event graph.

[0109] S108. Use a preset graph neural network to extract the relationship features in the target event graph, combine the attention weights to obtain the node feature sequence, and input the node feature sequence into the preset time series model to capture the dynamic evolution law, and obtain the event prediction graph according to the dynamic evolution law;

[0110] In this embodiment, the preset graph neural network first traverses and calculates the nodes and edges in the target event graph, identifies and extracts the causal relationship, temporal relationship, and semantic association between events, and then combines the attention weights adjusted by the cross-modal attention mechanism to integrate the node features and obtain the node feature sequence.

[0111] Subsequently, the obtained node feature sequence is input into a preset time series model. The preset time series model analyzes and models the time-varying data. By processing the node feature sequence, the dynamic evolution law of events in the time dimension is captured. Finally, based on the captured dynamic evolution law, an event prediction graph is generated. The event prediction graph not only contains information about the current event but also predicts and presents possible future events based on historical data and dynamic laws, enabling one to understand in advance the possible development directions and changes of events and prevent potential crises.

[0112] S109. Extract the causal chain in the event prediction graph, perform causal relationship reasoning based on the node feature sequence and the causal chain, and obtain the triggering condition and emotional factors after reasoning.

[0113] The event prediction graph contains rich event information and their prediction relationships. After sorting out the causal connections between the events in the event prediction graph and extracting the causal chain, which can clearly present the causal logic between events, causal relationship reasoning is performed based on the node feature sequence and the causal chain. In the reasoning process, the event attributes, emotional information contained in the node feature sequence and the logical relationship presented by the causal chain can be comprehensively considered.

[0114] Through causal relationship reasoning, the triggering condition and emotional factors are finally obtained. The triggering condition clarifies the prerequisite conditions and key factors for the occurrence of a certain event, while the emotional factors reveal how emotions drive the evolution of the event during the development of the event and the impact of different emotions on the event outcome.

[0115] Specifically, extract the causal chain in the event prediction graph, perform causal relationship reasoning based on the node feature sequence and the causal chain, and obtain the triggering condition and emotional factors, which are represented by the following formula:

[0116] ;

[0117] Among them, represents the triggering condition, represents the emotional factor, represents the triggering condition corresponding to the time step, the node feature sequence, represents the weight matrix parameter, represents the causal chain.

[0118] S110. Match the event prediction graph with the target event graph, and calculate the event analysis result by combining the node feature sequence, the triggering condition, and the emotional factor.

[0119] In this embodiment, the process of mutually matching the event prediction graph and the target event graph will compare the nodes, edges, and attribute information in the two graphs, and find the similarities and differences between the two graphs for matching and fusion. Subsequently, comprehensive analysis and calculation are carried out in combination with the node feature sequence, trigger condition, and emotional factor. In the process of calculation, various information resources are fully utilized, making the analysis results more comprehensive, in-depth, and accurate. Among them, the node feature sequence contains the detailed features and emotional information of the event, the trigger condition clarifies the premise for the occurrence of the event, and the emotional factor contains the impact of emotion on the event.

[0120] After the calculation is completed, the obtained event analysis results provide detailed decision-making basis for the urban governance department, enabling it to clearly understand the development dynamics, potential risks, and public emotional tendencies of the event, etc., facilitating the formulation of more scientific, reasonable, and effective response strategies, improving the efficiency and effect of urban governance, and better ensuring the safe and stable development of the city.

[0121] The event analysis method based on generation tasks and multi-modalities solves the limitation that the existing feature fusion technology based on deep learning only fuses the basic features of each modality by using multi-modal data, large model generation tasks, and cross-modal attention mechanisms, and realizes more accurate urban event analysis. First, the text, image, audio, and video features of the multi-modal data are extracted, and the cross-modal attention mechanism is used to dynamically allocate weights to form a multi-modal emotional relationship vector, which can more comprehensively and deeply explore the potential associations between data and improve the ability to capture event features. Then, by generating structured event information and constructing an event graph, and combining with a preset graph neural network to aggregate the multi-modal emotional relationship vector, an emotional feature vector is generated, closely integrating the emotional information of the event with the structure and relationship, and more accurately grasping the event trend and emotional dynamics.

[0122] After generating the sentiment feature vector, the contrastive learning loss value calculated by the preset contrastive loss function is used to optimize the sentiment feature vector and multimodal data. Then, the sentiment feature vector is input into the node as the node attribute of the event graph. The cross-modal attention mechanism is used to track the development and changes of events in real time, and the weight of the sentiment feature vector is dynamically adjusted to generate the target event graph. Then, the target event graph is processed by the preset graph neural network to extract the relationship features of the nodes. In the extraction process, the graph neural network can deeply explore the potential associations between nodes and provide reliable relationship information for subsequent analysis. After extracting the relationship features, the node feature sequence is obtained by combining the attention weights, and the node feature sequence is input into the preset time series model. After receiving the node feature sequence, the time series model can capture the dynamic evolution law of the event in the time series through iterative calculation, and generate the event prediction graph according to the dynamic evolution law, so as to predict the future development trend of the event. Finally, the event prediction graph is fused with the target event graph, and the event analysis result is obtained by combining the node feature sequence. The event analysis results obtained in this embodiment comprehensively consider the current status and future trends of the event, and can provide a more comprehensive and accurate basis for decision-making when faced with complex and changeable situations of urban events, which helps to improve the scientificity and timeliness of urban governance decisions, avoid missing the best time to deal with emergency events, and thus effectively control the scope of impact of the event.

[0123] See also Figure 2 , Figure 2 Another embodiment of an event analysis method based on generation tasks and multimodality provided by the present application includes:

[0124] S201, acquiring multimodal data, where the multimodal data includes text data, image data, and audio data;

[0125] In this embodiment, step S201 is similar to step S101 in the above embodiment and will not be described again here.

[0126] S202, setting a generation task based on a preset generation large model, inputting multimodal data into the generation task for training, and obtaining a structured description;

[0127] When setting a generation task based on the preset generation big model, first of all, according to the needs of urban situation monitoring and the characteristics of multimodal data, clarify the goal of the generation task, design input prompt words and output templates, and then preprocess the acquired multimodal data. The preprocessing includes removing noise, filtering irrelevant information and processing missing values. After the preprocessing is completed, the input prompt words and multimodal data are input into the generation task, and the preset generation big model in the generation task analyzes and trains the input prompt words and multimodal data.

[0128] During the training process, the preset generative large model will perform semantic understanding and analysis on text data to extract key information; perform feature recognition on image data to identify objects, human actions, etc. in the scene; and perform content parsing on audio data to determine the meaning represented by the sound. After training, a structured description is obtained according to the output template. For example, in the case of an urban fire incident, the preset generative large model in the generation task learns and trains on text, images of the fire scene, and rescue audio, and generates a structured description containing information such as the time of the fire, location, cause of the fire, development of the fire situation, and rescue operations. This structured description has higher accuracy and integrity and is more comprehensive than the description generated from single-modal data.

[0129] Specifically, designing input prompt words is to enable the preset generative large model to more accurately understand the goal of the generation task and generate structured event information. Examples of input prompt words are as follows:

[0130] "Please extract event information from the following text and generate structured event information, including event type, time, location, participants, victims, results, and causal relationships.";

[0131] The output template of the generation task needs to clearly define the event type and elements, and can use a JSON structure or natural language description:

[0132] An example of the JSON structure is as follows:

[0133] {

[0134] "Event type": "Urban safety event",

[0135] "Event": {

[0136] "Name": "Fire",

[0137] "Location": "Huangpu District, Shanghai",

[0138] "Time": "October 10, 2023",

[0139] "Involved persons": ["Residents", "Firefighters"],

[0140] "Actions": ["Fire started", "Rescue"],

[0141] "Cause": "Short circuit of the circuit"

[0142] }

[0143] }

[0144] An example of the natural language description is as follows:

[0145] “On October 10, 2023, a fire broke out in Huangpu District, Shanghai. The fire was caused by a short circuit in the wiring, and then the firefighters arrived to carry out the rescue.”

[0146] S203. Under the preset parameters of the currently preset generation large model, combine the multimodal data and the corresponding structured description into a sample pair, and calculate the target probability of the structured event information generated by the input multimodal data based on the sample pair;

[0147] Under the preset parameters of the currently preset generation large model, pair each set of obtained multimodal data with the structured description one by one into a sample pair, and then calculate the target probability of the structured description generated by the input multimodal data based on the sample pair. In the process of calculating the target probability, the preset generation large model will process and analyze the multimodal data according to the preset parameters, such as performing word vector conversion on the text, extracting and matching features of the image, and performing feature analysis on the audio, etc., and then calculate the percentage of the processed multimodal data to obtain the structured event information based on the sample pair. The calculated percentage result is the target probability of the multimodal data generating the structured event information.

[0148] The target probability can reflect the understanding and generation ability of the preset generation large model for different multimodal data, which is convenient for discovering the advantages and disadvantages of the preset generation large model, so as to improve and optimize the preset generation large model targeted.

[0149] S204. According to the target probability, combine the cross-entropy loss function to adjust the preset parameters, generate structured event information after adjustment, and perform relation extraction and semantic analysis on the structured event information to obtain the event relationship;

[0150] In this embodiment, according to the calculated target probability, combine the cross-entropy loss function to adjust the preset parameters, where the cross-entropy loss function is mainly used to measure the difference degree between the model prediction result and the real result. First, obtain the calculated target probability, calculate the loss value of generating the structured event information through the cross-entropy loss function, and according to the size of the loss value, use the optimization algorithm to perform backpropagation adjustment on the preset parameters of the preset generation large model. The optimization algorithm can use the stochastic gradient descent algorithm. Combining the cross-entropy loss function to adjust the preset parameters can continuously optimize the preset generation large model and make the structured event information generated by the preset generation large model more accurate and reliable.

[0151] Specifically, according to the target probability, combine the cross-entropy loss function to adjust the preset parameters, and generate structured event information after adjustment, which is represented by the following formula:

[0152] ;

[0153] Among them, Represents structured event information, Represents the quantity of multimodal data, Represents the target probability, Represents the preset parameters, Represents the th input multimodal data; Represents the th output structured description.

[0154] During the adjustment process, the preset generation large model will update the parameters in the direction of reducing the loss value, making the structured event information generated by the preset generation large model in the generation task closer to the true label. After the adjustment is completed, the structured event information is generated. Subsequently, relationship extraction between different elements in the event information is identified for the structured event information, and semantic analysis for deeply understanding the meaning of the event information is performed. Finally, the event relationship is obtained.

[0155] S205. Construct an event graph, use the structured event information as the nodes of the event graph, and connect the nodes according to the event relationship;

[0156] In this embodiment, step S205 is similar to step S103 in the foregoing embodiment, and will not be elaborated here.

[0157] S206. Use the preset language large model to extract the text semantics and text sentiment in the multimodal data, and obtain text features based on the text semantics and text sentiment;

[0158] In this embodiment, first, the text data in the multimodal data is input into the preset language large model. The language large model will perform lexical analysis on the text data, split the text data into individual words or phrases, and perform part-of-speech tagging on each word. Subsequently, syntactic analysis is performed to parse the grammatical structure of the sentence and determine the relationship between sentence components such as the subject, predicate, and object. Through lexical analysis and syntactic analysis, the preset language large model further understands the key information and theme content in the text data, and then extracts the text semantics.

[0159] While extracting the text semantics, the sentiment analysis module in the preset language large model is used to judge the sentiment tendency of the text data, determine whether the text data expresses positive, negative, or neutral sentiment. For example, it is judged whether the text is a praise for the urban safety measures or dissatisfaction with the urban safety incidents. After the judgment is completed, the judgment result is extracted as the text sentiment.

[0160] The extracted text semantics and text sentiment are subjected to feature encoding and conversion, and converted into numerical text features. For example, the text semantics is mapped to a point in a high-dimensional vector space, the text sentiment is represented by a specific numerical value and incorporated into the vector. Finally, the text features are obtained.

[0161] In this embodiment, the preset language large model is used to extract text semantics and text sentiment and obtain text features, which can deeply mine the rich information contained in the text part of the multimodal data. The extraction of text semantics helps to accurately understand people's descriptions and views on urban security events, and grasp the key points and core content of the events; text sentiment enables urban governance departments to understand the public's attitude towards the events.

[0162] S207. Use ResNet or Vision Transformer to extract the image scene and image expression in the multimodal data, and obtain image features based on the image scene and image expression;

[0163] In this embodiment, using ResNet or Vision Transformer to extract the image scene and image expression in the multimodal data can effectively extract key information from the image part of the multimodal data and obtain representative image features. Among them, the extraction of the image scene can intuitively display the environment and background where the urban security event occurs, and quickly understand the on-site situation of the event, while the extraction of the image expression can reflect people's emotional states during the event, providing important clues for understanding the public's reaction to the event.

[0164] When using ResNet for extraction, it is necessary to perform processing operations such as resizing the image and normalizing the pixel values on the image data in the multimodal data to make the image data meet the input requirements of the ResNet model. Then, the processed image data is input into the ResNet model, and the convolutional layer and pooling layer in the ResNet model are used to extract the image scene and image expression in the image data. Finally, the extracted image scene and image expression are converted into image features; when using Vision Transformer for extraction, the image can be directly segmented into multiple small pieces, and each small piece is respectively mapped into a vector sequence and input into the Vision Transformer model. The self-attention mechanism of the Vision Transformer model is used to process the vector sequence to capture the relationships between different small pieces, so as to identify the image scene and image expression. Finally, the extracted image scene and image expression are converted into image features. Converting the image scene and expression information into image features facilitates the fusion with the features of text data and audio data in the multimodal data, so as to more comprehensively analyze urban events.

[0165] During the extraction using ResNet or Vision Transformer, both the ResNet or Vision Transformer model will capture the detailed information in the image data, such as elements like buildings, roads, people, etc. in the image data, and then identify the scene depicted by the image data. By analyzing the facial features of people, it is possible to judge the expressions of the people in the image, such as fear, surprise, calmness, etc.

[0166] S208. Extract the audio intonation, audio volume, and audio rhythm in the multimodal data using the Wav2Vec audio model, and obtain audio features based on the audio intonation, audio volume, and audio rhythm.

[0167] In this embodiment, first, the audio data in the multimodal data is preprocessed. The preprocessing includes removing the noise interference in the audio, improving the clarity and quality of the audio, and ensuring that the format of the audio data meets the input requirements of the Wav2Vec audio model. Subsequently, the preprocessed audio data is input into the Wav2Vec audio model. The Wav2Vec audio model will perform frame-by-frame analysis on the preprocessed audio data, extract the audio intonation in the audio data through the convolutional layer, pooling layer, and fully connected layer of the Wav2Vec audio model, measure the volume of the audio data at the same time to understand the strength of the sound, and analyze the rhythm characteristics of the audio to obtain the audio rhythm. The audio intonation can convey the emotional state and attitude tendency of the speaker, the volume size can reflect the urgency or importance of the event, and the audio rhythm can represent the dynamic changes in the development of the event.

[0168] After the extraction is completed, based on the extracted audio intonation, audio volume, and audio rhythm, feature encoding is performed to convert them into numerical audio features. For example, for the audio at the scene of an urban security event, through the Wav2Vec audio model, the intonation of the crowd shouting, the volume of the alarm sound, and the rhythm in the rescue operation can be extracted, and these information can be converted into audio features.

[0169] This embodiment uses the Wav2Vec audio model to extract audio features, which can not only deeply understand urban events from the audio perspective, but also supplement the information that cannot be fully expressed by text and images, helping urban governance departments more accurately judge the atmosphere at the event scene, the emotional fluctuations of people, and the development trend of the event, and providing strong support for formulating more effective response strategies.

[0170] S209. Calculate the target weights of the text features, image features, and audio features using the cross-modal attention mechanism, and combine the target weights with the text features, image features, and audio features to obtain a multimodal sentiment relationship vector.

[0171] In this embodiment, the cross-modal attention mechanism is used to calculate the target weights of text features, image features, and audio features, which can effectively mine the internal connections between different modal data, and can dynamically allocate the proportion of different modal features in the overall information, highlighting key information and avoiding excessive interference of a certain modal data on the overall analysis due to incomplete or inaccurate information.

[0172] First, the extracted text features, image features, and audio features are input into the cross-modal attention mechanism model, which will compare and analyze the features of these three different modalities. For text features and image features, the cross-modal attention mechanism model will find the correlation points between text features and image features, such as whether the scene elements mentioned in the text features are reflected in the image features, so as to measure the correlation between text features and image features; for text features and audio features, the cross-modal attention mechanism model will analyze the connection between the emotions expressed by the text features and the audio intonation, volume, etc. in the audio features; for image features and audio features, the cross-modal attention mechanism model will explore whether the image scene in the image features matches the audio rhythm and sound atmosphere of the audio features. Through these comparative analyses, the cross-modal attention mechanism model can calculate the target weights of each modal feature in the overall information expression, and finally perform weighted fusion on the target weights and the corresponding text features, image features, and audio features to obtain a multi-modal sentiment relationship vector.

[0173] The multi-modal sentiment relationship vector fully combines the feature advantages of the three modalities of text, image, and audio, is no longer limited to the information of a single modality, contains more comprehensive and rich sentiment and relationship information, can more accurately reflect the real situation and internal sentiment tendency of the event, helps urban governance departments to deeply analyze the event from multiple dimensions, better grasp the public's sentiment attitude and the potential relationship between events, so as to formulate more practical response measures and further improve the decision-making quality and effect of urban safety management.

[0174] S210. Aggregate the multi-modal sentiment relationship vector based on a preset graph neural network, and combine it with the event graph to obtain a sentiment feature vector;

[0175] S211. Construct positive sample pairs and negative sample pairs according to the sentiment feature vector and multi-modal data. Based on the positive sample pairs and negative sample pairs, use a preset contrast function to calculate the contrast learning loss value, and optimize the alignment of the sentiment feature vector and multi-modal data through the contrast learning loss value to obtain the target sentiment feature;

[0176] S212. Transmit the sentiment feature vector to the nodes of the event graph, and use the cross-modal attention mechanism to adjust the attention weights of the sentiment feature vector to generate the target event graph;

[0177] S213. Use a preset graph neural network to extract relationship features in the target event graph, combine the attention weights to obtain a node feature sequence, and input the node feature sequence into a preset time series model to capture the dynamic evolution law, and obtain an event prediction graph according to the dynamic evolution law;

[0178] S214. Extract the causal chain in the event prediction graph, perform causal relationship reasoning based on the node feature sequence and the causal chain, and obtain the trigger condition and emotional factors after reasoning;

[0179] S215. Match the event prediction graph with the target event graph, and combine the node feature sequence, trigger condition, and emotional factors to calculate and obtain the event analysis result.

[0180] In this embodiment, steps S210 to S215 are similar to steps S105 to S110 in the foregoing embodiment, and will not be elaborated here.

[0181] S216. Convert the data structure of the event analysis result into a target data structure, and use a preset visualization tool to generate a visual analysis result for the converted event analysis result.

[0182] In this embodiment, first, a detailed analysis is performed on the data structure of the event analysis result to clarify the existing data organization form, data elements, and the association relationship between the elements of the event analysis result. According to the actual requirements of the target data structure, data conversion rules are designed. For example, if the original data structure is a hierarchical structure and the target data structure is a key-value pair structure that is convenient for quick query, then the original data needs to be reorganized and mapped in the form of key-value pairs. Subsequently, the data in the event analysis result is converted into the target data structure according to the data conversion rules. After the conversion is completed, a preset visualization tool is used, and appropriate visualization chart types are selected according to different dimensions and requirements of the event analysis result. For example, a bar chart is used to compare the occurrence frequencies of different events, a line chart is used to show the development trend of events, and a map is used to intuitively present the distribution of event occurrence locations, etc. Finally, a visual analysis result is generated.

[0183] Generating a visual analysis result using a preset visualization tool improves the readability and comprehensibility of the event analysis result. The visual charts can present complex and abstract event analysis data in an intuitive graphical form, enabling quick understanding of the key information, development trend, and the relationship between different factors of the event.

[0184] The following provides a detailed description of the event analysis system based on generation tasks and multi-modalities provided by the present application. Please refer to Figure 3 , Figure 3 This is another embodiment of the event analysis system based on generation tasks and multi-modalities provided by the present application. The system includes:

[0185] An acquisition unit 301, configured to acquire multimodal data, where the multimodal data includes text data, image data, and audio data;

[0186] A first generation unit 302, configured to generate a task based on a preset generation large model, input the multimodal data into the generation task, generate structured event information, and perform relation extraction and semantic analysis on the structured event information to obtain event relations;

[0187] A construction unit 303, configured to construct an event graph, use the structured event information as nodes of the event graph, and connect the nodes according to the event relations;

[0188] A fusion unit 304, configured to extract text features, image features, and audio features of the multimodal data, and use a cross-modal attention mechanism to fuse the text features, image features, and audio features to obtain a multimodal sentiment relation vector;

[0189] An aggregation unit 305, configured to aggregate the multimodal sentiment relation vector based on a preset graph neural network, and combine it with the event graph to obtain a sentiment feature vector;

[0190] A first calculation unit 306, configured to construct positive sample pairs and negative sample pairs based on the sentiment feature vector and the multimodal data, calculate a contrastive learning loss value using a preset contrastive function based on the positive sample pairs and the negative sample pairs, and perform alignment optimization on the sentiment feature vector and the multimodal data through the contrastive learning loss value to obtain a target sentiment feature;

[0191] A second generation unit 307, configured to transmit the sentiment feature vector to the nodes of the event graph, and use a cross-modal attention mechanism to adjust the attention weights of the sentiment feature vector to generate a target event graph;

[0192] A capture unit 308, configured to use a preset graph neural network to extract relation features in the target event graph, combine the attention weights to obtain a node feature sequence, input the node feature sequence into a preset time series model to capture the dynamic evolution law, and obtain an event prediction graph according to the dynamic evolution law;

[0193] An inference unit 309, configured to extract the causal chain in the event prediction graph, perform causal relation inference based on the node feature sequence and the causal chain, and obtain a trigger condition and an emotional factor after the inference;

[0194] A second calculation unit 310, configured to match the event prediction graph with the target event graph, and calculate an event analysis result in combination with the node feature sequence, the trigger condition, and the emotional factor.

[0195] Optionally, the fusion unit 304 is specifically configured to:

[0196] Use a pre - set language large - model to extract the text semantics and text sentiment in multi - modal data, and obtain text features based on the text semantics and text sentiment;

[0197] Use ResNet or Vision Transformer to extract the image scene and image expression in multi - modal data, and obtain image features based on the image scene and image expression;

[0198] Use the Wav2Vec audio model to extract the audio intonation, audio volume, and audio rhythm in multi - modal data, and obtain audio features based on the audio intonation, audio volume, and audio rhythm;

[0199] Use a cross - modal attention mechanism to calculate the target weights of the text features, image features, and audio features, and combine the target weights with the text features, image features, and audio features to obtain a multi - modal sentiment relationship vector.

[0200] Optionally, the first generation unit 302 is specifically used for:

[0201] Set a generation task based on a pre - set generation large - model, input multi - modal data into the generation task for training, and obtain a structured description;

[0202] Under the pre - set parameters of the current pre - set generation large - model, combine the multi - modal data and the corresponding structured description into a sample pair, and calculate the target probability of the structured event information generated by the input multi - modal data based on the sample pair;

[0203] Adjust the pre - set parameters according to the target probability in combination with the cross - entropy loss function, and generate structured event information after adjustment.

[0204] Optionally, adjusting the pre - set parameters according to the target probability in combination with the cross - entropy loss function, and generating structured event information after adjustment, is represented by the following formula:

[0205] ;

[0206] Wherein, represents the structured event information, represents the number of multi - modal data, represents the target probability, represents the pre - set parameters, represents the th input multi - modal data; represents the th output structured description.

[0207] Optionally, it further includes a third generation unit 311, which is specifically used for:

[0208] Convert the data structure of the event analysis result into the target data structure, and use a preset visualization tool to generate a visual analysis result from the converted event analysis result.

[0209] Optionally, extract the causal chain in the event prediction graph, perform causal relationship reasoning based on the node feature sequence and the causal chain, and obtain the triggering condition and emotional factor after reasoning, which are represented by the following formula:

[0210] ;

[0211] Among them, represents the triggering condition, represents the emotional factor, represents the corresponding triggering condition in the time step, represents the node feature sequence, represents the weight matrix parameter, represents the causal chain.

[0212] Optionally, the contrast learning loss value is represented by the following formula:

[0213] ;

[0214] Among them, represents the contrast loss value, represents the emotional feature vector of the positive sample pair, represents the emotional feature vector of the negative sample pair, represents the adjustable parameter, represents the number of multimodal data, represents the similarity between the emotional feature vectors of the positive sample pair and the negative sample pair, represents the traversal index variable.

[0215] This application also provides an event analysis device based on generation tasks and multimodality. Please refer to Figure 4 , Figure 4 which is an embodiment of the event analysis device based on generation tasks and multimodality provided by this application. The device includes:

[0216] a processor 401, a memory 402, an input / output unit 403, and a bus 404;

[0217] The processor 401 is connected to the memory 402, the input / output unit 403, and the bus 404;

[0218] The memory 402 stores a program, and the processor 401 calls the program to execute any of the above methods.

[0219] This application also relates to a computer-readable storage medium with a program stored thereon, characterized in that when the program runs on a computer, it causes the computer to execute any of the above methods.

[0220] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0221] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0222] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0223] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0224] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.

Claims

1. An event analysis method based on generation tasks and multimodality, characterized in that: The method comprises: Acquiring multimodal data, wherein the multimodal data includes text data, image data, and audio data; A generation task is set based on a preset generation model, the multimodal data is input into the generation task, structured event information is generated, and relationship extraction and semantic analysis are performed on the structured event information to obtain event relationships; Constructing an event graph, taking the structured event information as nodes of the event graph, and connecting the nodes according to the event relationships; Extracting text features, image features, and audio features of the multimodal data, and fusing the text features, the image features, and the audio features using a cross-modal attention mechanism to obtain a multimodal sentiment relationship vector; Aggregating the multimodal sentiment relationship vector based on a preset graph neural network, and obtaining a sentiment feature vector in combination with the event graph; A positive sample pair and a negative sample pair are constructed according to the emotional feature vector and the multimodal data. Based on the positive sample pair and the negative sample pair, a contrastive learning loss value is calculated using a preset contrast function. The emotional feature vector and the multimodal data are aligned and optimized by the contrastive learning loss value to obtain a target emotional feature. The contrastive learning loss value is expressed by the following formula: ; in, represents the contrast loss value, represents the sentiment feature vector of the positive sample pair, Represents the sentiment feature vector of the negative sample pair, Represents adjustable parameters, represents the amount of the multimodal data, represents the similarity between the sentiment feature vectors of the positive sample pair and the negative sample pair, Indicates traversal indicator variables; The emotion feature vector is transmitted to the node of the event graph, and the attention weight of the emotion feature vector is adjusted by using the cross-modal attention mechanism to generate a target event graph; The preset graph neural network is used to extract the relationship features in the target event graph, and the node feature sequence is obtained by combining the attention weight, and the node feature sequence is input into the preset time series model to capture the dynamic evolution law, and the event prediction graph is obtained according to the dynamic evolution law; The causal chain in the event prediction graph is extracted, and causal relationship reasoning is performed based on the node feature sequence and the causal chain. After reasoning, the trigger conditions and emotional factors are obtained, which are expressed by the following formula: ; in, represents the trigger condition, Indicates the emotional factors, express The trigger condition corresponding to the time step, represents the node feature sequence, represents the weight matrix parameters, represents the causal chain; The event prediction graph is matched with the target event graph, and the event analysis result is calculated by combining the node feature sequence, the trigger condition and the emotional factor.

2. The method according to claim 1, characterized in that The extracting of text features, image features and audio features of the multimodal data, and fusing the text features, the image features and the audio features using a cross-modal attention mechanism to obtain a multimodal sentiment relationship vector includes: Extracting text semantics and text sentiment from the multimodal data using a preset language macro model, and obtaining text features based on the text semantics and the text sentiment; Extracting image scenes and image expressions from the multimodal data using ResNet or Vision Transformer, and obtaining image features based on the image scenes and the image expressions; Extracting audio intonation, audio volume, and audio rhythm from the multimodal data using a Wav2Vec audio model, and obtaining audio features based on the audio intonation, the audio volume, and the audio rhythm; The target weights of the text features, the image features, and the audio features are calculated using a cross-modal attention mechanism, and the target weights are combined with the text features, the image features, and the audio features to obtain a multimodal sentiment relationship vector.

3. The method according to claim 1, characterized in that The step of setting a generation task based on a preset generation large model, inputting the multimodal data into the generation task, and generating structured event information includes: Setting a generation task based on a preset generation model, inputting the multimodal data into the generation task for training, and obtaining a structured description; Under the preset parameters of the preset large model, the multimodal data and the corresponding structured description are combined into sample pairs, and the target probability of the structured event information generated by the input multimodal data is calculated based on the sample pairs; According to the target probability, the preset parameters are adjusted in combination with a cross entropy loss function, and structured event information is generated after the adjustment.

4. The method according to claim 3, characterized in that The preset parameters are adjusted according to the target probability in combination with the cross entropy loss function, and structured event information is generated after the adjustment, which is expressed by the following formula: ; in, Represents structured event information, represents the amount of the multimodal data, represents the target probability, represents the preset parameters, Indicates The multimodal data of the input; Indicates The structured description of the output.

5. The method according to claim 1, characterized in that: After matching the event prediction graph with the target event graph, combining the node feature sequence, the trigger condition and the emotional factor to calculate the event analysis result, the method further includes: The data structure of the event analysis result is converted into a target data structure, and a preset visualization tool is used to generate a visualization analysis result from the converted event analysis result.

6. An event analysis system based on generation tasks and multimodality, characterized in that: The system comprises: An acquisition unit, configured to acquire multimodal data, wherein the multimodal data includes text data, image data, and audio data; A first generation unit is used to set a generation task based on a preset generation model, input the multimodal data into the generation task, generate structured event information, and perform relationship extraction and semantic analysis on the structured event information to obtain event relations; A construction unit, used to construct an event graph, use the structured event information as nodes of the event graph, and connect the nodes according to the event relationship; A fusion unit, used to extract text features, image features and audio features of the multimodal data, and fuse the text features, the image features and the audio features using a cross-modal attention mechanism to obtain a multimodal sentiment relationship vector; An aggregation unit, used to aggregate the multimodal sentiment relationship vector based on a preset graph neural network, and obtain a sentiment feature vector in combination with the event graph; The first calculation unit is used to construct a positive sample pair and a negative sample pair according to the emotional feature vector and the multimodal data, and based on the positive sample pair and the negative sample pair, use a preset contrast function to calculate a contrastive learning loss value, and align and optimize the emotional feature vector and the multimodal data by using the contrastive learning loss value to obtain a target emotional feature, wherein the contrastive learning loss value is expressed by the following formula: ; in, represents the contrastive learning loss value, represents the sentiment feature vector of the positive sample pair, Represents the sentiment feature vector of the negative sample pair, Represents adjustable parameters, represents the amount of the multimodal data, represents the similarity between the sentiment feature vectors of the positive sample pair and the negative sample pair, Indicates traversal indicator variables; A second generating unit is used to transmit the emotion feature vector to the node of the event graph, adjust the attention weight of the emotion feature vector by using the cross-modal attention mechanism, and generate a target event graph; A capture unit, used to extract the relationship features in the target event graph using the preset graph neural network, obtain a node feature sequence in combination with the attention weight, and input the node feature sequence into a preset time series model to capture the dynamic evolution law, and obtain an event prediction graph according to the dynamic evolution law; The reasoning unit is used to extract the causal chain in the event prediction graph, perform causal relationship reasoning based on the node feature sequence and the causal chain, and obtain the triggering conditions and emotional factors after reasoning, which are expressed by the following formula: ; in, represents the trigger condition, Indicates the emotional factors, express The trigger condition corresponding to the time step, represents the node feature sequence, represents the weight matrix parameters, represents the causal chain; The second calculation unit is used to match the event prediction graph with the target event graph, and calculate the event analysis result by combining the node feature sequence, the trigger condition and the emotional factor.

7. An event analysis device based on generation tasks and multi-modality, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Construction method and system of multi-modal affair graph and readable storage medium

    CN114020936A

  • Multimodal sentiment classification method and system for modal sequence perception of global audio feature enhancement

    CN116189039A