Social event detection method based on self-supervised multi-modal fusion
Through the self-supervised multimodal fusion social event detection method, using twin neural networks and hierarchical clustering algorithms for structural entropy discrimination, the dependence on single-modal data and detection problems in the open environment in traditional methods are solved, and efficient and accurate social event detection is achieved to adapt to the multimodal changes of social media.
Patent Information
- Application Number
- CN202510611635.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-05
AI Technical Summary
The existing social event detection methods rely on single-modal data and cannot fully capture the full picture and details of events. It is difficult to adapt to the multimodal change characteristics of social media data and the detection of new events in an open environment. It requires manual labeling and predefined event categories, resulting in insufficient ability to handle large-scale, diverse and dynamically evolved social events.
The social event detection method of self-supervised multimodal fusion is adopted. Through data acquisition, preprocessing, self-supervised training and event detection, a twin neural network (SNN) is used for comparison learning and knowledge distillation, combined with a multimodal large language model enhancement module and event encoder, the deep fusion of text and image features is achieved, and a hierarchical clustering algorithm for structural entropy discrimination is used for cluster detection.
Without manual annotation and predefined event categories, social events can be detected efficiently and accurately in the open world, adapting to complex and changeable social media data, improving the accuracy and completeness of event feature representation, simplifying the model training process, and improving detection efficiency and adaptability.
Smart Images

Figure CN120429670A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of social event detection, and in particular relates to a social event detection method based on self-supervised multimodal fusion. Background Art
[0002] In recent years, various online platforms (such as Google News, Twitter, Flickr, etc.) have developed rapidly, and users can easily share text, images, videos and other content anytime and anywhere. Therefore, the data characteristics of social message streams have gradually shown a multimodal trend. Traditional social event detection (SED) methods only support unimodal data (mainly text or images). However, unimodal data often cannot fully capture the overall picture and details of the event, which can easily lead to information loss or misjudgment, and is increasingly unsuitable for the data stream characteristics of current social media platforms. Therefore, multimodal social event detection (MSED) is very important. How to design an effective MSED framework is the challenge that the present invention needs to solve.
[0003] Currently, due to the limitations of the model framework, the existing MSED and SED methods have strict requirements on data quality and format, and require additional manpower for manual labeling. For example, supervised methods require manual labeling of events in the original data for model training. Some graph neural network (GNN)-based methods require relevant attributes such as user mentions, topic tags, user IDs, and keywords to aggregate the features of neighbor nodes. However, in the real world, these attributes are often partially missing or require manual labeling. These limitations result in the current methods being insufficient in handling large-scale, diverse, and dynamically evolving social events. Therefore, how to free the model from strict data format requirements to adapt to the changing characteristics of social media is a challenge that the present invention needs to solve.
[0004] Furthermore, while some research is exploring MSED, a drawback of these studies is that they often simplify the MSED task. Most studies treat MSED as a supervised classification task. These approaches are applicable only to event detection in closed environments where the total number of event types is known and predefined. However, in the real world, new events occur daily, and predefined event labels often do not align with real-world characteristics, making these supervised classification methods ineffective in open environments. Although recent research has attempted to enable classifiers to recognize labels not seen in training samples at test time to discover new events in open environments, these approaches can only roughly classify new events into the same category and fail to provide more fine-grained classification or recognition. Thanks to research on unimodal SED of text, it is more reasonable and effective to treat MSED as a clustering task. However, due to the complexity of multimodal data, data fusion and representation suitable for clustering are significant challenges. To our knowledge, no MSED method currently treats event detection as a clustering task. The challenge addressed by this paper is to formalize the MSED task as a clustering process that meets the requirements of open-world detection. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a social event detection method based on self-supervised multimodal fusion, focusing on the field of open-world multimodal social event detection, aiming to overcome the limitations of traditional methods and use innovative technologies to achieve more efficient and accurate event detection.
[0006] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a social event detection method based on self-supervised multimodal fusion, comprising the steps of:
[0007] Data collection: collecting multimodal data including text and images from various social media platforms; data preprocessing, including data cleaning, text segmentation, and image normalization;
[0008] Self-supervised training uses SNN as the training network; the multimodal large language model enhancement module and event encoder I constitute the teacher model, which receives preprocessed text and images as input and encodes the input into event features F t ; Event encoder II and predictor constitute the student model, which maps event features to a new vector space represented as F s , which forms a positive sample pair for contrastive learning with the output of the teacher model; the goal of this network is to make the output of the student model gradually approach the output of the teacher model;
[0009] Event detection,The original text and image are input into the student model to obtain the latent space representation of the event, and then clustering is performed to obtain the social event detection results.
[0010] Furthermore, the data preprocessing includes:
[0011] Data cleaning: Clean the collected data to identify and remove duplicate data; use data validation rules and error correction mechanisms to identify and remove erroneous data; and remove irrelevant content.
[0012] Text segmentation processing, using the word segmentation tools and algorithms in natural language processing to divide continuous text into vocabulary units;
[0013] Image standardization processing, through image editing and conversion technology, adjusts image parameters such as size and resolution.
[0014] Furthermore, in the multimodal large language model enhancement module, the event is enhanced from three aspects: event type E type 、Event Theme E theme and image description E caption ;
[0015] The augmented text will be represented as the concatenation of the original text and the augmented text:
[0016]
[0017] Among them, O text Represents the original text, Represents a join operation.
[0018] Furthermore, in the self-supervised learning process, the visual question answering (VQA) model is used for processing, including:
[0019] (1) Constructing positive sample pairs for contrastive learning: The model needs to learn to distinguish between similar and dissimilar data pairs to obtain effective feature representations; the visual question answering (VQA) model can mine the potential knowledge in multimodal data to determine which data corresponds to the same event and thus construct positive sample pairs;
[0020] (2) Mining multimodal latent knowledge: When a text-image pair is input, the VQA model performs word segmentation on the text and extracts keywords. At the same time, it performs scene classification and object recognition on the image to extract key visual information. Then, through a rule-based and machine learning inference engine combined with a predefined knowledge graph, it infers relevant information.
[0021] (3) Obtain rich text representation: Input questions about the image into the VQA model; the VQA model generates corresponding text descriptions based on the image content and the visual and language knowledge learned through pre-training.
[0022] (4) Assisted generation of enhanced data: The VQA model generates corresponding enhanced information based on clues in text and images and its knowledge reserves;
[0023] The original text is concatenated with the enhanced information generated by VQA to obtain the enhanced text.
[0024] Furthermore, for the text set T = {t1, t2, ..., t n} and image set I={i1,i2,...,i m}, each text t j ∈T and image i k ∈I is input into the VQA model; the VQA model contains a visual feature extractor and a text encoder, which extract the visual features of the image and the semantic features of the text respectively; then, a matching network is used to calculate the matching score between the two; if the matching score exceeds the preset threshold τ, it is determined that the text and the image describe the same event, and (t j ,i k ) as a positive sample pair.
[0025] Furthermore, the event encoder I and event encoder II adopt the same structure, and the event encoder is used to achieve deep fusion of image and text features to represent event introduction;
[0026] First, the event encoder uses the pre-trained visual model ViT and the pre-trained language model LM;
[0027] For ViT, using the visual encoder in the CLIP model, ViT aligned with the text will facilitate cross-modal interaction; freezing ViT to prevent its parameters from updating, thereby achieving the goal of integrating image features into text features in the event encoder; the visual feature is represented as: F vision =ViT(img), img is the image data;
[0028] For LM, the SBERT model is used and average pooling is applied, that is, the last hidden state of the words in the event is averaged to obtain its embedding; the text feature is represented as: F text =LM(text), where text is text data;
[0029] Subsequently, the event encoder uses a cross-attention mechanism to fuse image features, where the query Q, key K, and value V are F text 、F vision and F vision ; The attention score output S is defined as:
[0030]
[0031] Among them, d kWith F text The dimensions are the same, d k is the scaling factor;
[0032] The fused features are defined as S and F vision The product of: F fused =SF vision ;
[0033] Next, the fused features are further characterized by a multi-layer perceptron MLP to obtain the final event feature F event ; The MLP of the event encoder consists of three simple linear layers, and the final F event is encoded as a vector.
[0034] Furthermore, the core goal of the predictor is to make accurate cluster predictions of events based on the features output by the event encoder, while ensuring the stability of the model during training;
[0035] The predictor adopts the architecture of multi-layer perceptron (MLP). The predictor consists of multiple fully connected layers and activation functions. It gradually transforms and maps the input features and finally outputs the predicted results of the event.
[0036] Furthermore, the predictor includes two linear layers forming an MLP, which maps the event features output by the event encoder into a new vector space, and forms a positive sample pair for contrastive learning with the output of the teacher model;
[0037] The first linear layer receives the feature F1 output by the event encoder and passes it through the weight matrix W h1 and the bias vector b h1 Perform linear transformation to obtain the intermediate feature H h1 ; The expression is: H h1 =W h1 F1+b h1 ; This step performs preliminary processing on the input features, adjusting their dimensions and feature combinations;
[0038] The second linear layer H h1 For further transformation, use the weight matrix W h2 and the bias vector b h2 , and get the final output vector F s , the expression is: F s =W h2 ·H h1 +b h2 , this output vector F s and the event feature F output by the teacher model t Constitute positive sample pairs for calculating contrast loss and driving the model to learn more effective event feature representations.
[0039] Furthermore, during event detection, a latent space representation of the event is obtained, including the following steps:
[0040] The input of the latent space mapping network is assumed to be the fusion feature F i , after multiple layers of linear transformation and nonlinear activation function ReLU, it is mapped into the latent space;
[0041] Assume that the weight matrix of the mapping network is W1,W2,...,W m , the bias vectors are b1,b2,...,b m , then the latent space representation Z i The calculation formula is:
[0042] H1=ReLU(W1F i +b1)
[0043] H2=ReLU(W2H1+b2) ...
[0045] Z i =W m H m-1 +b m
[0046] Among them, H1 is the intermediate feature representation after the first layer transformation, H2 is the intermediate feature representation after the second layer transformation, and H m-1 is the output after the penultimate layer transformation, Z i is the feature representation in the latent space.
[0047] Furthermore, a hierarchical clustering algorithm based on structural entropy discrimination is used for clustering during event detection;
[0048] The hierarchical clustering algorithm based on structural entropy discrimination maps each layer of the hierarchical clustering tree to a coding tree with a height of 2 to calculate the two-dimensional structural entropy and automatically selects the optimal number of clusters n by minimizing the structural entropy.
[0049] The hierarchical clustering algorithm based on structural entropy discrimination includes the following steps:
[0050] (1) Build graph structure: Initialize graph Where V consists of N leaf nodes, and the edge set E is initially empty; for each leaf node v i , calculate the cosine similarity between it and all other nodes one by one; based on the calculation results, each leaf node v i Connect to the K nodes with the highest similarity to it and build a graph The edge set E of the node preliminarily establishes the connection relationship between the nodes, providing the basic data structure for the subsequent structural entropy calculation and cluster analysis;
[0051] (2) Mapping and calculation of structural entropy: First, initialize the structural entropy set SE to an empty set and set the counter n to N; then, enter the loop iteration process, as long as n>0, perform the following operations: map the corresponding level of the current clustering tree to a coding tree T with a height of 2 n The coding tree is the key data structure for calculating structural entropy. Through this mapping method, we can deeply analyze the clustering situation from the perspective of graph structure.
[0052] For each coding tree T n , according to the formula:
[0053]
[0054] Calculate its two-dimensional structural entropy 2DSE;
[0055] Among them, n j is the number of nodes in the partition, is the weighted degree of the nodes in the partition, V j is the sum of the weighted degrees of all nodes in the partition, is the sum of the weights of the cut edges in the partition, w is the sum of the weighted degrees of all nodes, m represents the number of partitions the graph is divided into, j is the partition index, and i is the node index within the partition;
[0056] 2DSE is an important indicator for measuring the stability of node partitioning and clustering quality. The smaller its value is, the more stable the current clustering structure is and the better the clustering effect is.
[0057] Each time a coding tree T is calculated n 2D SE value se n After that, add it to the structural entropy set SE and reduce the counter n by 1; continue this process until n = 0, at which point the calculation of the structural entropy of the coding tree corresponding to each layer of the clustering tree is completed;
[0058] (3) Determine the optimal number of clusters and the result: After obtaining a series of 2D SE values, find the index n corresponding to the maximum value from the set SE best , maximum value index n best Corresponding to the minimum 2D SE value, the minimum 2D SE value represents the optimal clustering method; based on this index n best , get the corresponding clustering results This clustering result is used as the final event detection result.
[0059] The beneficial effects of adopting this technical solution are:
[0060] The present invention designs a novel simple twin network model for open-world multimodal social event detection. The present invention carefully designs a twin neural network (SNN) that can perform contrastive learning and knowledge distillation simultaneously. The present invention learns in a self-supervised manner, which eliminates the need for manual participation in data labeling during the learning process, and avoids the trouble of pre-defining events required by traditional methods, greatly simplifying the model training process. During the training process, the present invention uses a multimodal large language model to enhance the original events to obtain higher-level knowledge. In the detection stage, the present invention designs a hierarchical clustering algorithm based on structural entropy SE discrimination to realize event detection, and the model of the present invention only requires original text and original images as input, without any additional attributes.
[0061] In order to simply and efficiently encode event features into the latent space in the open world, the present invention uses a student model to complete feature encoding. It is worth noting that the student model of the present invention only requires input of raw text and images, without any manual annotation or additional information, which will greatly improve the efficiency of practical applications. Since the model of the present invention is trained in a self-supervised manner and is not restricted by any labels, it can be easily used for event feature representation in the open world, even for newly emerging events. Events will be encoded into the latent space. Then, the present invention uses an unsupervised clustering method to cluster the events in the latent space to achieve event detection.
[0062] The present invention solves the problem of multimodal data processing: social media data presents multimodal characteristics, and traditional social event detection methods rely on unimodal data, making it difficult to fully capture event information. The present invention uses a multimodal large language model enhancement module to enhance the original data from three aspects: event type, theme, and image subtitles to obtain a high-order representation of the event. At the same time, combined with technologies such as twin neural networks (SNN), it deeply integrates text and image features, effectively processes multimodal data, and comprehensively and accurately captures the complete information of the event, making up for the shortcomings of unimodal data and improving the accuracy and completeness of event feature representation.
[0063] The present invention breaks away from the limitations of data labeling and pre-definition: existing methods have strict requirements on data quality and format, relying on manual labeling and pre-defining event categories, which face many difficulties in practical applications. The present invention adopts a self-supervised learning approach in SNN, using contrastive learning and knowledge distillation technology, eliminating the need for manual data labeling and pre-defining event categories. During the learning process, by controlling the direction of gradient propagation, contrastive learning and knowledge distillation can be carried out simultaneously, greatly simplifying the model training process, reducing data processing costs, and improving the adaptability and versatility of the model, enabling it to better cope with complex and changing social events in the open world.
[0064] The present invention realizes efficient event detection in the open world: Events in the open world are complex and diverse, and new events are constantly emerging. The existing multimodal social event detection (MSED) method simply regards the task as a supervised classification task and cannot effectively handle new events. The present invention formalizes the MSED task as a clustering task and designs a hierarchical clustering algorithm based on structural entropy discrimination. The algorithm maps each layer of the hierarchical clustering tree to a coding tree, calculates 2D SE, and automatically selects the optimal number of clusters. Without predefining the number of events, it can efficiently detect social events in an open world environment, improve the accuracy and efficiency of detection, and provide strong support for social management, fake news detection, public safety and other fields.
[0065] This invention improves model performance and stability: By carefully designing the SNN structure and training strategy, this invention ensures model stability during contrastive learning and knowledge distillation. The predictor plays a key role in stabilizing model training and avoiding model crashes during training.
[0066] This invention meets practical application needs: It operates efficiently in the face of massive amounts of social media data. Relying solely on raw text and images as input, it can rapidly complete detection tasks, meeting the need for real-time detection and providing timely and accurate reference information for downstream tasks. In terms of visualization, it accurately categorizes events and intuitively presents event clustering results, facilitating user understanding and analysis, and better serving practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 Schematic diagram of a social event detection method based on self-supervised multimodal fusion according to the present invention;
[0068] Figure 2 This is a diagram of the event encoder framework in an embodiment of the present invention;
[0069] Figure 3 It is a hierarchical clustering algorithm framework based on structural entropy discrimination in an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.
[0071] In this embodiment, see Figure 1 As shown, the present invention proposes a social event detection method based on self-supervised multimodal fusion, comprising the steps of:
[0072] Data collection: Collect multimodal data containing text and images from various social media platforms, such as Twitter and Weibo. (These data come from a rich and diverse source, covering information about social events across different fields and types, providing ample material for subsequent analysis and research. By collecting user content posted on these platforms, we can obtain multimodal information closely related to real-world events, laying the data foundation for multimodal social event detection.) Data preprocessing, including data cleaning, text segmentation, and image normalization.
[0073] Self-supervised training uses SNN as the training network; the multimodal large language model enhancement module and event encoder I constitute the teacher model, which receives preprocessed text and images as input and encodes the input into event features F t ; Event encoder II and predictor constitute the student model, which maps event features to a new vector space represented as F s , which forms a positive sample pair for contrastive learning with the output of the teacher model; the goal of this network is to make the output of the student model gradually approach the output of the teacher model;
[0074] Event detection,The original text and image are input into the student model to obtain the latent space representation of the event, and then clustering is performed to obtain the social event detection results.
[0075] As an optimization solution of the above embodiment, the data preprocessing includes:
[0076] Data cleaning: Clean the collected data, identify and remove duplicate data, prevent duplicate information from interfering with model training, ensure data uniqueness, and improve data quality; use data validation rules and error correction mechanisms to identify and eliminate erroneous data to ensure data accuracy; remove irrelevant content such as special symbols and invalid labels. These redundant information may interfere with the model's extraction of valid information. The cleaned data is more concise and standardized, which is conducive to subsequent model processing.
[0077] Text segmentation processing uses word segmentation tools and algorithms in natural language processing to divide continuous text into vocabulary units. This processing method can convert text into a format suitable for model input, making it easier for the model to analyze and understand text information, thereby extracting the semantic features contained therein.
[0078] Image normalization involves adjusting image parameters such as size and resolution through image editing and conversion techniques. For example, images of varying sizes can be resized to the standard size required for model training, while the resolution can be optimized to meet the training requirements. This ensures consistency when inputting images from different sources and sizes into the model, preventing poor model training results due to image format differences, thereby improving the effectiveness and stability of model training.
[0079] As an optimization solution of the above embodiment, in the multimodal large language model enhancement module, the event is enhanced from three aspects: event type E type 、Event Theme E theme and image description E caption .
[0080] For event type E type , it can roughly understand events such as natural disasters, financial crises, terrorist attacks, etc. The present invention uses the following prompts to obtain E type : "Post content: [text]. Please read the post content and image and tell me what type of event this is?".
[0081] For event theme E theme , which can provide a more detailed understanding of the event, including specific themes or topics within the event type, such as "hurricane consequences", "market collapse impact" or "terrorist attack details", which will help the present invention to judge the specific event. The present invention uses the following prompts to obtain E theme : "This is an image and text of a social event: [text]. Please summarize the specific theme of the event from the image and text!".
[0082] For the image description E caption , which can generate descriptive captions for images, helping to understand the visual content of the post. It is worth noting that E caption Only as visual auxiliary information, more important visual features will be fused through the cross-modal attention mechanism (introduced in the event encoder). This paper uses the following hints to obtain E caption : "Please provide a brief description of the image!".
[0083] Finally, the augmented text will be represented as the concatenation of the original text and the augmented text:
[0084]
[0085] Among them, O text Represents the original text, Represents a join operation.
[0086] As an optimization solution for the above embodiment, a visual question answering (VQA) model is used in the self-supervised learning process to perform the following processing:
[0087] (1) Constructing positive sample pairs for contrastive learning: The model needs to learn to distinguish between similar and dissimilar data pairs to obtain effective feature representations; the visual question answering (VQA) model can be used to mine the potential knowledge in multimodal data to determine which data corresponds to the same event and then construct positive sample pairs.
[0088] (2) Mining Multimodal Latent Knowledge: The VQA model can deeply analyze the relationship between text descriptions and image content, and infer key information such as event type and topic from multimodal data. When a text-image pair is input, the VQA model performs word segmentation on the text and extracts keywords. At the same time, it performs scene classification and object recognition on the image to extract key visual information of the image. Then, through a rule-based and machine learning inference engine combined with a predefined knowledge graph, it infers relevant information.
[0089] Taking event type inference as an example, if keywords such as "earthquake" and "tsunami" appear in text t, and image i shows corresponding disaster scenes, such as collapsed houses and seawater flooding, the VQA model infers that the event type is a natural disaster. Let the VQA model be VQA(t,i), and its output is the event type e type , that is, e type =VQA(t,i), where t is text and i is image.
[0090] To infer the event topic, the VQA model considers both the specific description in the text and the detailed information in the image. For example, if the text mentions "rescue operation" and the image shows rescuers working in the rubble, the VQA model will infer that the event topic is "earthquake rescue."
[0091] (3) Obtaining Rich Text Representations: VQA mines knowledge in a question-and-answer format, transforming image and text information into textual descriptions with greater semantic depth, thereby enhancing the expressive power of textual features. Questions about an image are fed to the VQA model, such as "Please describe the scene in the image" or "What are the main characters in the image doing?" The VQA model generates a corresponding textual description based on the image content and the visual and linguistic knowledge learned through pre-training.
[0092] Assume that the original text is t original , the description generated by VQA is t vqa , then the enriched text is represented as in Represents a text concatenation operation. For example, if the original text is "This is a disaster" and the VQA-generated description is "The picture shows a city destroyed by an earthquake, with ruins and rescue workers everywhere", the enriched text will be "This is a disaster. The picture shows a city destroyed by an earthquake, with ruins and rescue workers everywhere."
[0093] (4) Assisted generation of enhanced data: With specific prompts, the VQA model can conduct in-depth analysis based on the input text and image content and output relevant enhanced information. These prompts can be "supplement the time and place of the event" or "describe the impact of the event." The VQA model generates corresponding enhanced information based on the clues in the text and image and combines its knowledge reserves.
[0094] For example, if the input text mentions "a flood", the image shows a flooded city street, and the prompt is "Add the time and place where the event occurred", the VQA model may output "The flood occurred in [specific city] at [specific time]".
[0095] The original text is concatenated with the enhanced information generated by VQA to obtain the enhanced text.
[0096] For a text set T = {t1, t2, ..., t n} and image set I={i1,i2,...,i m}, each text t j ∈T and image i k ∈I is input into the VQA model; the VQA model contains a visual feature extractor and a text encoder, which extract the visual features of the image and the semantic features of the text respectively; then, a matching network is used to calculate the matching score between the two; if the matching score exceeds the preset threshold τ, it is determined that the text and the image describe the same event, and (t j ,i k ) as a positive sample pair.
[0097] The judgment process can be expressed by the following formula:
[0098] if MatchScore(VQA(t j ,i k ))>
[0099] τthen(t j ,i k )is a positive pair.
[0100] As an optimization solution of the above embodiment, the event encoder I and event encoder II adopt the same structure, and the event encoder is used to achieve deep fusion of image and text features to represent event introduction; Figure 2 As shown:
[0101] First, the event encoder uses the pre-trained vision model ViT and the pre-trained language model LM.
[0102] For ViT, using the visual encoder in the CLIP model, ViT aligned with the text will facilitate cross-modal interaction; freezing ViT to prevent its parameters from updating, thereby achieving the goal of integrating image features into text features in the event encoder; the visual feature is represented as: F vision =ViT(img), img is image data.
[0103] For LM, the SBERT model is used and average pooling is applied, that is, the last hidden state of the words in the event is averaged to obtain its embedding; the text feature is represented as: F text =LM(text), where text is text data.
[0104] Subsequently, the event encoder fuses the image features using a cross-attention mechanism, where the query Q, key K, and value V are F text 、F vision and F vision ; The attention score output S is defined as:
[0105]
[0106] Among them, d k With F text The dimensions are the same, s k is the scaling factor, which is a hyperparameter;
[0107] The fused features are defined as S and F vision The product of: F fused =SF vision .
[0108] Next, the fused features are further characterized by a multi-layer perceptron MLP to obtain the final event feature F event ; The MLP of the event encoder consists of three simple linear layers, and the final F event is encoded as a vector.
[0109] In the event encoder, MLP is used to further extract and transform the fused image and text features to obtain the final event feature representation. After obtaining image and text features through the visual transformer (ViT) and language model (LM), and fusing these features using the cross attention mechanism, the fused feature F is obtained. fused Although the fused features at this point already contain multimodal information, they still need further processing to better represent events.
[0110] The MLP in the event encoder consists of three simple linear layers. The first linear layer receives the fused feature F fused , and perform a preliminary linear transformation on it, and map the input features to a new vector space through the weight matrix W1 and the bias vector b1. The formula is:
[0111] H1=W1·F fused +b1
[0112] This step adjusts the dimension and distribution of features in preparation for subsequent processing.
[0113] Next, the second linear layer transforms H1 again, also using the weight matrix W2 and bias vector b2, namely:
[0114] H2=W2·H1+b2
[0115] This layer further extracts and combines features to discover potential patterns in the data.
[0116] Finally, the third linear layer maps H2 to the final event feature F event , encoded as a 384-dimensional vector, the expression is:
[0117] F event =W3·H2+b3
[0118] After this series of linear transformations, MLP transforms the fused features into more representative and discriminative event features, providing effective data representation for subsequent contrastive learning and event detection.
[0119] As an optimization solution to the above-mentioned embodiment, this network incorporates a predictor to effectively utilize the features extracted by the event encoder in the multimodal social event detection task and further enhance the model's ability to predict events. The core goal of the predictor is to accurately cluster and predict events based on the features output by the event encoder while ensuring the stability of the model during training.
[0120] The predictor uses a multi-layer perceptron (MLP) architecture, which boasts a simple structure and powerful nonlinear mapping capabilities, effectively processing the complex features output by the event encoder. Specifically, the predictor consists of multiple fully connected layers and activation functions, gradually transforming and mapping the input features to ultimately output the event prediction.
[0121] Let the feature vector output by the event encoder be z, and the input of the predictor be z. The predictor processes z through a series of linear transformations and nonlinear activation functions, and finally outputs the prediction vector p. The specific formula is as follows:
[0122] p=h(z)=σ(W n ...σ(W2σ(W1z+b1)+b2)...+b n )
[0123] Among them, W i and b i are the weight matrix and bias vector of the i-th fully connected layer, and σ is the activation function.
[0124] Preferably, the predictor comprises two linear layers forming an MLP, which maps the event features output by the event encoder into a new vector space, and forms a positive sample pair for contrastive learning with the output of the teacher model;
[0125] The first linear layer receives the feature F1 output by the event encoder and passes it through the weight matrix W h1 and the bias vector b h1 Perform linear transformation to obtain the intermediate feature H h1 ; The expression is: H h1 =W h1 F1+b h1 ; This step performs preliminary processing on the input features, adjusting their dimensions and feature combinations;
[0126] The second linear layer H h1 For further transformation, use the weight matrix W h2 and the bias vector b h2 , and get the final output vector F s , the expression is: F s =W h2 ·H h1 +b n2 , this output vector F s and the event feature F output by the teacher model t Constitute positive sample pairs for calculating contrast loss and driving the model to learn more effective event feature representations.
[0127] As an optimization solution of the above embodiment, the event features obtained by the student model are still high-dimensional. In order to further reduce the complexity of the data and mine potential patterns, the present invention maps them to a latent space.
[0128] When detecting an event, we obtain the potential space representation of the event, including the following steps:
[0129] The input of the latent space mapping network is assumed to be the fusion feature F i , after multiple layers of linear transformation and nonlinear activation function ReLU, it is mapped into the latent space;
[0130] Assume that the weight matrix of the mapping network is W1,W2,...,W m, the bias vectors are b1,b2,...,b m , then the latent space representation Z i The calculation formula is:
[0131] H1=ReLU(W1F i +b1)
[0132] H2=ReLU(W2H1+b2) ...
[0134] Z i =W m H m-1 +b m
[0135] Among them, H1 is the intermediate feature representation after the first layer transformation, H2 is the intermediate feature representation after the second layer transformation, and H m-1 is the output after the penultimate layer transformation, Z i is the feature representation in the latent space.
[0136] Through this mapping, the present invention compresses high-dimensional fusion features into a low-dimensional latent space while retaining the key information of the data.
[0137] A hierarchical clustering algorithm based on structural entropy discrimination is used for clustering during event detection. The hierarchical clustering algorithm based on structural entropy discrimination maps each layer of the hierarchical clustering tree to a coding tree with a height of 2 to calculate the two-dimensional structural entropy and automatically selects the optimal number of clusters n by minimizing the structural entropy.
[0138] In order to detect events in latent space without predefining the total number of events, the present invention proposes a hierarchical clustering algorithm based on structural entropy discrimination. Hierarchical clustering is a traditional machine learning method that can construct a cluster tree (also known as a dendrogram) by merging smaller clusters into larger clusters. However, a disadvantage of hierarchical clustering is that the number of clusters needs to be predefined, which is often impractical in practical applications because the number of clusters is usually unknown. Therefore, the present invention attempts to map each layer of the hierarchical clustering tree to a coding tree with a height of 2 to calculate the two-dimensional structural entropy, and automatically select the optimal number of clusters n by minimizing the structural entropy.
[0139] Hierarchical clustering algorithms based on structural entropy discrimination, such as Figure 3 As shown, the steps include:
[0140] (1) Build graph structure: Initialize graph Where V consists of N leaf nodes, and the edge set E is initially empty; for each leaf node v i, calculate the cosine similarity between it and all other nodes one by one; cosine similarity is a widely used similarity metric that can effectively measure the similarity between vectors. In this algorithm, by calculating the cosine similarity between the feature vectors corresponding to the leaf nodes, the similarity between different event messages can be accurately quantified. Based on the calculation results, each leaf node v i Connect to the K nodes with the highest similarity to it and build a graph The edge set E of the node preliminarily establishes the connection relationship between the nodes, providing the basic data structure for the subsequent structural entropy calculation and cluster analysis;
[0141] (2) Mapping and calculation of structural entropy: First, initialize the structural entropy set SE to an empty set and set the counter n to N; then, enter the loop iteration process, as long as n>0, perform the following operations: map the corresponding level of the current clustering tree to a coding tree T with a height of 2 n The coding tree is the key data structure for calculating structural entropy. Through this mapping method, we can deeply analyze the clustering situation from the perspective of graph structure.
[0142] For each coding tree T n , according to the formula:
[0143]
[0144] Calculate its two-dimensional structural entropy 2DSE;
[0145] Among them, n j is the number of nodes in the partition, is the weighted degree of the nodes in the partition, V j is the sum of the weighted degrees of all nodes in the partition, is the sum of the weights of the cut edges in the partition, w is the sum of the weighted degrees of all nodes, m represents the number of partitions the graph is divided into, j is the partition index, and i is the node index within the partition;
[0146] 2DSE is an important indicator for measuring the stability of node partitioning and clustering quality. The smaller its value is, the more stable the current clustering structure is and the better the clustering effect is.
[0147] Each time a coding tree T is calculated n 2D SE value se n After that, add it to the structural entropy set SE and reduce the counter n by 1; continue this process until n = 0, at which point the calculation of the structural entropy of the coding tree corresponding to each layer of the clustering tree is completed;
[0148] (4) Determine the optimal number of clusters and the result: After obtaining a series of 2D SE values, find the index n corresponding to the maximum value from the set SE best, maximum value index n best Corresponding to the minimum 2D SE value, the minimum 2D SE value represents the optimal clustering method; based on this index n best , get the corresponding clustering results This clustering result is used as the final event detection result.
[0149] This strategy of automatically selecting the optimal number of clusters based on structural entropy successfully avoids the difficulty of pre-defining the number of clusters in traditional hierarchical clustering algorithms, allowing the algorithm to more flexibly and accurately adapt to situations where the number of events is unknown in actual application scenarios, thereby efficiently and accurately detecting events in the latent space.
[0150] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A social event detection method based on self-supervised multimodal fusion, characterized in that: Including steps: Data collection: collecting multimodal data including text and images from various social media platforms; data preprocessing, including data cleaning, text segmentation, and image normalization; Self-supervised training uses SNN as the training network; the multimodal large language model enhancement module and event encoder I constitute the teacher model, which receives preprocessed text and images as input and encodes the input into event features F t ; Event encoder II and predictor constitute the student model, which maps event features to a new vector space represented as F s , which forms a positive sample pair for contrastive learning with the output of the teacher model; the goal of this network is to make the output of the student model gradually approach the output of the teacher model; Event detection,The original text and image are input into the student model to obtain the latent space representation of the event, and then clustering is performed to obtain the social event detection results.
2. The social event detection method based on self-supervised multimodal fusion according to claim 1 is characterized in that: The data preprocessing includes: Data cleaning: Clean the collected data to identify and remove duplicate data; use data validation rules and error correction mechanisms to identify and remove erroneous data; and remove irrelevant content. Text segmentation processing, using the word segmentation tools and algorithms in natural language processing to divide continuous text into vocabulary units; Image standardization processing, through image editing and conversion technology, adjusts image parameters such as size and resolution.
3. The social event detection method based on self-supervised multimodal fusion according to claim 1 is characterized in that: In the multimodal large language model enhancement module, events are enhanced from three aspects: event type E type 、Event Theme E theme and image description E caption ; The augmented text will be represented as the concatenation of the original text and the augmented text: Among them, O text Represents the original text, Represents a join operation.
4. The social event detection method based on self-supervised multimodal fusion according to claim 3 is characterized in that: In the self-supervised learning process, the visual question answering (VQA) model is used for processing, including: (1) Constructing positive sample pairs for contrastive learning: The model needs to learn to distinguish between similar and dissimilar data pairs to obtain effective feature representations; the visual question answering (VQA) model can mine the potential knowledge in multimodal data to determine which data corresponds to the same event and thus construct positive sample pairs; (2) Mining multimodal latent knowledge: When a text-image pair is input, the VQA model performs word segmentation on the text and extracts keywords. At the same time, it performs scene classification and object recognition on the image to extract key visual information. Then, through a rule-based and machine learning inference engine combined with a predefined knowledge graph, it infers relevant information. (3) Obtaining rich text representation: Input questions about the image to the VQA model; the VQA model generates a corresponding text description based on the image content and the visual and language knowledge learned through pre-training; (4) Assisted generation of enhanced data: The VQA model generates corresponding enhanced information based on clues in text and images and its knowledge reserves; The original text is concatenated with the enhanced information generated by VQA to obtain the enhanced text.
5. The social event detection method based on self-supervised multimodal fusion according to claim 4 is characterized in that: For a text set T = {t1, t2, ..., t n } and image set I={i1,i2,...,i m }, each text t j ∈T and image i k ∈I is input into the VQA model; the VQA model contains a visual feature extractor and a text encoder, which extract the visual features of the image and the semantic features of the text respectively; then, a matching network is used to calculate the matching score between the two; if the matching score exceeds the preset threshold τ, it is determined that the text and the image describe the same event, and (t j ,i k ) as a positive sample pair.
6. The social event detection method based on self-supervised multimodal fusion according to claim 1, characterized in that: The event encoder I and event encoder II adopt the same structure, and the event encoder is used to achieve deep fusion of image and text features to represent event introduction; First, the event encoder uses the pre-trained visual model ViT and the pre-trained language model LM; For ViT, using the visual encoder in the CLIP model, ViT aligned with the text will facilitate cross-modal interaction; freezing ViT to prevent its parameters from updating, thereby achieving the goal of the event encoder integrating image features into text features; The visual feature is represented as: F vision =ViT(img), img is the image data; For LM, the SBERT model is used and average pooling is applied, that is, the last hidden states of the words in the event are averaged to obtain their embeddings; The text feature is represented as: F text =LM(text), where text is text data; Subsequently, the event encoder fuses the image features using a cross-attention mechanism, where the query Q, key K, and value V are F text 、F vision and F vision ; The attention score output S is defined as: Among them, d k With F text The dimensions are the same, d k is the scaling factor; The fused features are defined as S and F vision The product of: F fused =SF vision ; Next, the fused features are further characterized by a multi-layer perceptron MLP to obtain the final event feature F event ; The MLP of the event encoder consists of three simple linear layers, and the final F event is encoded as a vector.
7. The social event detection method based on self-supervised multimodal fusion according to claim 1, characterized in that: The core goal of the predictor is to make accurate cluster predictions of events based on the features output by the event encoder while ensuring the stability of the model during training. The predictor adopts the architecture of multi-layer perceptron (MLP). The predictor consists of multiple fully connected layers and activation functions. It gradually transforms and maps the input features and finally outputs the predicted results of the event.
8. The social event detection method based on self-supervised multimodal fusion according to claim 7 is characterized in that: The predictor consists of two linear layers forming an MLP, which maps the event features output by the event encoder into a new vector space, which forms a positive sample pair for contrastive learning with the output of the teacher model; The first linear layer receives the feature F1 output by the event encoder and passes it through the weight matrix W h1 and the bias vector b h1 Perform linear transformation to obtain the intermediate feature H h1 ; The expression is: H h1 =W h1 F1+b h1 ; This step performs preliminary processing on the input features, adjusting their dimensions and feature combinations; The second linear layer H h1 For further transformation, use the weight matrix W h2 and the bias vector b h2 , and get the final output vector F s , the expression is: F s =W h2 ·H h1 +b h2 , this output vector F s and the event feature F output by the teacher model t Constitute positive sample pairs for calculating contrast loss and driving the model to learn more effective event feature representations.
9. The social event detection method based on self-supervised multimodal fusion according to claim 1, characterized in that: When detecting an event, we obtain the potential space representation of the event, including the following steps: The input of the latent space mapping network is assumed to be the fusion feature F i , after multiple layers of linear transformation and nonlinear activation function ReLU, it is mapped into the latent space; Assume that the weight matrix of the mapping network is W1,W2,...,W m , the bias vectors are b1,b2,...,b m , then the latent space representation Z i The calculation formula is: Among them, H1 is the intermediate feature representation after the first layer transformation, H2 is the intermediate feature representation after the second layer transformation, and H m-1 is the output after the penultimate layer transformation, Z i is the feature representation in the latent space.
10. The social event detection method based on self-supervised multimodal fusion according to claim 9, characterized in that: In event detection, a hierarchical clustering algorithm based on structural entropy discrimination is used for clustering; The hierarchical clustering algorithm based on structural entropy discrimination maps each layer of the hierarchical clustering tree to a coding tree with a height of 2 to calculate the two-dimensional structural entropy and automatically selects the optimal number of clusters n by minimizing the structural entropy. The hierarchical clustering algorithm based on structural entropy discrimination includes the following steps: (1) Build graph structure: Initialize graph Where V consists of N leaf nodes, and the edge set E is initially empty; for each leaf node v i , calculate the cosine similarity between it and all other nodes one by one; Based on the calculation results, each leaf node v i Connect to the K nodes with the highest similarity to it and build a graph The edge set E of the node preliminarily establishes the connection relationship between the nodes, providing the basic data structure for the subsequent structural entropy calculation and cluster analysis; (2) Mapping and calculation of structural entropy: First, initialize the structural entropy set SE to an empty set and set the counter n to N; then, enter the loop iteration process, as long as n>0, perform the following operations: map the corresponding level of the current clustering tree to a coding tree T with a height of 2 n The coding tree is the key data structure for calculating structural entropy. Through this mapping method, we can deeply analyze the clustering situation from the perspective of graph structure. For each coding tree T n , according to the formula: Calculate its two-dimensional structural entropy 2DSE; Among them, n j is the number of nodes in the partition, is the weighted degree of the nodes in the partition, V j is the sum of the weighted degrees of all nodes in the partition, is the sum of the weights of the cut edges in the partition, w is the sum of the weighted degrees of all nodes, m represents the number of partitions the graph is divided into, j is the partition index, and i is the node index within the partition; 2DSE is an important indicator for measuring the stability of node partitioning and clustering quality. The smaller its value is, the more stable the current clustering structure is and the better the clustering effect is. Each time a coding tree T is calculated n 2D SE value se n After that, add it to the structural entropy set SE and reduce the counter n by 1; continue this process until n = 0, at which point the calculation of the structural entropy of the coding tree corresponding to each layer of the clustering tree is completed; (3) Determine the optimal number of clusters and the result: After obtaining a series of 2D SE values, find the index n corresponding to the maximum value from the set SE best , maximum value index n best Corresponding to the minimum 2D SE value, the minimum 2D SE value represents the optimal clustering method; based on this index n best , get the corresponding clustering results This clustering result is used as the final event detection result.
Citation Information
Cited By
Social governance work order dispatching method based on emergency degree evaluation, medium and equipment
CN120822923A