Incomplete multi-modal crisis event detection method based on memory pool and modal perception expert system

By using a memory pool and a modal perception expert system, the problem of incomplete modalities in multimodal crisis event detection was solved, achieving high-precision information recovery and modal fusion, and improving the accuracy and flexibility of detection.

CN120974443AActive Publication Date: 2025-11-18BEIJING UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511507810.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing technologies cannot accurately recover and complete missing modal information in multimodal crisis event detection, resulting in insufficient detection accuracy and timeliness.

Method used

We employ a method based on memory pools and modality-aware expert systems. By constructing cross-modal memory pools, deep feature interaction, and modality-aware expert systems, we achieve missing modality completion and feature extraction. We combine graph-to-text and text-to-graph large-scale models to generate high-quality completion data, and perform modality fusion through multi-head attention mechanism and lightweight routing network.

Benefits of technology

It improves the accuracy and robustness of multimodal crisis event detection, can accurately recover information in the case of missing modalities, enhances the accuracy and flexibility of detection, adapts to deep interactions between different modalities, and enhances the model's adaptability and real-time response capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974443A_ABST
    Figure CN120974443A_ABST
Patent Text Reader

Abstract

The invention discloses an incomplete multi-mode crisis event detection method based on a memory pool and a mode perception expert system, and relates to the field of machine learning, and the method specifically comprises the steps: building a prompt strongly related to an event instance based on the memory pool, and complementing a missing mode through a large model; a CNN encoder, a ResNet encoder and a CLIP encoder are utilized to construct a double-flow encoder, and text and visual feature extraction is achieved; capturing fine-grained event information and clues; the distribution difference between a complemented sample and a complete sample is identified based on a modal perception expert system, and the quality of fusion representation is improved; and inputting the mixed features into the classification head to complete a crisis event detection task. According to the invention, through the prompt construction module based on the memory pool, a large model is introduced to realize careful completion of a missing mode, and the distribution difference between a completion sample and a complete sample is accurately identified based on a mode perception expert system, so that a multi-mode data missing scene can be effectively coped with, and the accuracy and timeliness of crisis event detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of machine learning, in particular to an incomplete multi-modal crisis event detection method based on a memory pool and a modal perception expert system. BACKGROUND

[0002] In multi-modal crisis event detection, data from different modalities, including images and text, are usually relied on for comprehensive analysis and identification. These modalities can provide rich scene information, for example, identifying fire scenes through images or analyzing emergency reports through text. However, multi-modal data in the real world often has incomplete situations, especially in actual crisis events, due to data loss, transmission delay or platform limitations, only single modality data is often available.

[0003] For example, in some sudden crisis events, image data may be lost or unavailable due to transmission problems, while only text data can provide a preliminary description of the event. Conversely, in some cases, image data may be available, but lacks corresponding textual descriptions. This situation of incomplete modalities leads to the loss of cross-modal information, affecting the accuracy and timeliness of event detection.

[0004] To address this challenge, how to recover and complete the missing modalities when the missing modal information exists has become a key technical problem to improve the effectiveness of multi-modal crisis event detection systems. In existing technologies, many methods rely on simple interpolation techniques or generate missing modalities based on existing data, but these methods often fail to accurately capture the specific context or characteristics of crisis events, resulting in inaccurate recovered modal information that cannot meet the high-precision detection requirements.

[0005] Therefore, how to generate high-quality completed data for missing modalities through intelligent modal completion mechanisms, and effectively combine the completed modalities with existing modalities, is an important direction to improve multi-modal crisis event detection technology. SUMMARY

[0006] In order to achieve accurate and timely crisis event detection in the presence of partial modality loss, the present application proposes an incomplete multi-modal crisis event detection method based on a memory pool and a modal perception expert system, including the following steps: The incomplete multi-modal crisis event detection method based on a memory pool and a modal perception expert system includes the following steps: Missing modal completion: based on the text and image pairs in the training set, a cross-modal memory pool is constructed, in which the text and image are respectively encoded to extract feature representations and stored in the memory pool in the form of vectors. For samples with missing text modal, the corresponding image features are used as query vectors to conduct cross-modal retrieval based on cosine similarity in the text feature set of the memory pool. For samples with missing image modal, the text features are used as query vectors to conduct cross-modal retrieval based on cosine similarity in the image feature set of the memory pool. After retrieval, the P pairs of memory samples with the highest cosine similarity are selected, and the Q keywords with the highest frequency are extracted from the text data in these samples. Subsequently, for samples with missing text modal, the original image features, the P pairs of text data retrieved and the extracted keywords are used to construct prompt information, which is input into the image-to-text large model to generate completed text. For samples with missing image modal, the original text features, the P pairs of text data retrieved and the keywords are used to construct prompt, which is input into the text-to-image large model to complete the image, P=3, Q=3. Text and visual feature extraction: TEXTCNN encoder and CLIP text encoder are used to extract text features, ResNet and CLIP visual encoder are used to extract image features, and the features are spliced in dimension, and image text contrast loss is used to reduce the feature difference between the completed features and the real features. Deep feature interaction: the multi-head attention mechanism based on transformer is used to realize the deep interaction of text and image features, extract fine-grained feature information, and mine subtle event clues in images and text. Gated expert modal perception: different experts are used to process different modal combinations to identify the modal feature differences between samples. Crisis event classification: the classification head is used to process the mixed features to identify the event category to which the current sample belongs.

[0007] As a further scheme of the present application, in the missing modal completion step, if the missing modal is text, the image-to-text large model is used to synthesize the missing description, and if the missing modal is image, the text-to-image large model generates a visually consistent image through reverse diffusion under the guidance of the prompt.

[0008] As a further scheme of the present application, in the text and visual feature extraction step, for text features, the TEXTCNN encoder and the pre-trained CLIP text encoder are used to extract modal-specific features and modal-consistent features: For image features, ResNET architecture and pre-trained CLIP image encoder are used to extract modal-specific features and modal-consistent features: in, and These represent modality-specific and modality-consistent encoders of text, respectively. and These represent modality-specific and modality-consistent encoders of an image, respectively. and This indicates modality-specific and modality-consistent features of text. and Represents modality-specific and modality-consistent features of an image; Concatenate modality-specific and modality-consistent features as output features for text and image modalities: in, or , This indicates a splicing operation.

[0009] As a further aspect of the present invention, the image-text contrast loss method involved is as follows: Where s is a learnable temperature parameter used to control the smoothness of the contrast learning; Represents the similarity moments from text features to image features; A similarity matrix representing image features to text features; Represents the transpose of the image matrix; This indicates that the temperature parameter has been exponentialized. The contrastive loss for each sample is as follows: Where B is the size of each batch of training samples; For text-to-image contrast loss; For image-to-text contrast loss; Indicates the index of the current sample; Indicates the index of the candidate matching sample.

[0010] The final comparison loss is: in, To compare the hyperparameters of the loss weights.

[0011] As a further aspect of the present invention, in the deep feature interaction step, a hybrid serial-parallel encoder is used to perform fine-grained feature extraction and capture subtle event cues. The multi-head self-attention mechanism based on the transformer is used to process the text and image features respectively, to enhance the representation of each modality, to calculate the self-attention of each attention head, and to independently enhance the text and image features: wherein, , and are the query, key and value embeddings of the i-th head, is a learnable parameter matrix, represents the input feature dimension, is the output dimension, and then the self-attention of each attention head is calculated: wherein, represents the output matrix of the i-th attention head, represents the concatenation matrix of the attention outputs of the m heads, is the final output matrix representing the multi-head attention, , represents the number of heads, represents the concatenation operation, represents the output linear transformation matrix; Subsequently, after the FFN network composed of two fully connected layers and a ReLU activation function, the enhanced features of the modal are obtained: wherein, represents the multi-head attention operation; The enhanced text and image features are fused through the cross-modal fusion module to generate the mixed features: .

[0012] wherein, represents the image enhanced feature, is the text enhanced feature, and are the mixed features of the image and the text.

[0013] ​​As a further scheme of the present application, the combination of different modalities is processed by different experts respectively, comprising the following steps: Adopting a lightweight routing network According to the semantic content and the detected modal configuration, a prediction adaptive routing score is configured; when the text modal is completed, the crisis perception routing network first calculates the score of each routing: Among them, respectively represent the scores of three specific modal experts , , The score of the general expert is selected from the general expert set according to their routing scores, and is recorded as The routing score of the specific modal expert The normalized routing weight is calculated by the softmax function: Among them, represents the weight assigned to the specific modal expert and the selected general expert; Each expert processes the input features respectively: Among them, is the output of the specific modal expert, is the output of the th general expert.

[0014] The final output feature is obtained by the following formula: .

[0015] As a further scheme of the present application, according to the mixed features, the classification of crisis events is carried out: Among them, is a full connection layer, and each layer has an activation function except the last layer; A standard cross-entropy loss function is adopted to optimize the classification performance of crisis event detection: Among them, N is the total number of samples, C is the number of categories, is the true label, is the predicted probability of category c; The total loss function is as follows: .​​

[0016] Compared with the prior art, the present application has the advantages that: in response to the problem of incomplete modal in multi-modal crisis event detection, the present application exhibits higher precision and robustness. First, based on the memory pool, the prompt related to the event instance is constructed, and the large model is combined for modal completion, so that even in the case of missing a certain modal, the system can restore the missing information and ensure the accuracy of the detection result. Traditional methods often rely on simple interpolation or rule filling, which is difficult to ensure the context consistency and event perception ability of the completed information. Second, by using CNN, ResNet and CLIP encoder to construct a double-flow encoder and deep feature interaction, subtle event clues in images and texts can be carefully mined, so as to extract more rich feature information. Finally, the modal perception expert system can flexibly dispatch different experts for processing according to the routing score of the modal specific and general experts, so as to realize more accurate modal fusion and information completion. This dynamic routing-based design enables the model to interact more deeply between different modalities, while avoiding the limitations of fixed expert structure, greatly improving the adaptability and accuracy of the system. In addition, based on the classification head, the mixed features are processed, which can effectively fuse the information of different modalities, thereby improving the recognition accuracy of complex event classes. In summary, the present application not only solves the challenge brought by incomplete modal, but also improves the real-time response ability, accuracy and flexibility of the model, which has important application value in crisis event detection. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 The flowchart of the method of the present application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described below in a clear and complete manner. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] The incomplete multi-modal crisis event detection method based on memory pool and modal perception expert system proposed by the present application realizes accurate and timely multi-modal crisis event detection while accurately completing the missing modal information through the following implementation steps. The implementation method includes missing modal completion, text and image feature extraction, deep feature interaction, gated expert modal perception and crisis event classification steps, each of which can be adjusted and optimized according to task requirements. The implementation method of each step is as follows.

[0020] (1) Step one: missing modal completion This step is the starting point of model training, ensuring that missing modal information can be effectively completed. This step includes three parts: building a memory pool, extracting similar text and keywords, and generating and issuing missing modalities: A. Building a memory pool: The present application builds a high-quality text-visual memory pool from a complete training set, in which text and visual features are encoded using a CLIP encoder.

[0021] B. Extracting similar text and keywords: Based on the cosine similarity of the features, the current modal information is retrieved in the memory pool to obtain the most similar 3 texts. For the extracted text, KeyBERT is used to extract keywords from the text, and the top three words with the highest word frequency are selected.

[0022] C. Missing modality generation: If the missing modality is text, the InstructBLIP model is used to synthesize the missing description. If the missing modality is an image, the text-to-image model (such as Stable Diffusion) generates a visually consistent image through back diffusion under the guidance of the prompt.

[0023] (2) Step two: text and image feature extraction: This step fully utilizes the visual-linguistic relationship to extract comprehensive feature information, and the implementation method is as follows: A. Text feature extraction: For text features, the TEXCNN architecture and pre-trained CLIP text encoder are used to extract modality-specific and modality-consistent features: B. Image feature extraction: For image features, the ResNET architecture and pre-trained CLIP image encoder are used to extract modality-specific and modality-consistent features: wherein, and represent the modality-specific and modality-consistent encoders of the text, and represent the modality-specific and modality-consistent encoders of the image, and represent the modality-specific and modality-consistent features of the text, and represent the modality-specific and modality-consistent features of the image; Concatenate the modality-specific and modality-consistent features as the output features of the text and image modalities: in, or , Indicates a splicing operation; C. Comparison of losses: To further narrow the gap between the modality-specific feature space and the modality-invariant feature space, an image-text contrast loss was employed. Where s is a learnable temperature parameter used to control the smoothness of the contrast learning; Represents the similarity moments from text features to image features; A similarity matrix representing image features to text features; Represents the transpose of the image matrix; This indicates that the temperature parameter has been exponentialized. The contrastive loss for each sample is as follows: Where B is the size of each batch of training samples; For text-to-image contrast loss; For image-to-text contrast loss; Indicates the index of the current sample; Indicates the index of the candidate matching sample.

[0024] The final comparison loss is: in, To compare the hyperparameters of the loss weights.

[0025] (3) Step 3: Deep feature interaction This method enables deep interaction between text and image features, allowing for detailed mining of subtle event clues within images and text, thereby extracting richer feature information. The implementation method is as follows: A transformer-based multi-head self-attention mechanism is employed to process text and image features separately, enhancing the representation of each modality. Self-attention is computed for each attention head, independently enhancing text and image features. in, , and They are the first The query of size, key and value embedding, For learnable parameter matrix, denotes the input feature dimension, is the output dimension, and then the self-attention of each attention head is calculated, wherein, denotes the output matrix of the th attention head, denotes the splicing matrix of the attention output of the m heads, is the final output matrix of the multi-head attention, , denotes the number of heads, denotes the splicing operation, denotes the output linear transformation matrix; Subsequently, after adopting the FFN network composed of two fully connected layers and ReLU activation function, the enhanced features of the modal can be obtained, wherein, denotes the multi-head attention operation; The enhanced text and image features are fused through the cross-modal fusion module to generate mixed features: .

[0026] wherein, denotes the image enhanced feature, is the text enhanced feature, and are the mixed features of the image and the text.

[0027] (4) Step four: gated expert modal perception By assigning a dedicated expert model for different modal combinations, the feature differences between each modal can be accurately identified, thereby avoiding conflicts or loss during information fusion. The implementation method is as follows: A lightweight routing network is adopted to predict adaptive routing scores according to semantic content and detected modal configurations. Taking the complete text modal as an example, the crisis perception routing network first calculates the score of each route, wherein, denote the scores of the three specific modal experts , , , denotes the general expert Score From the general set of experts, we select the top k experts based on their routing scores, denoted as [k experts]. Combining specific modality experts Routing score We calculate the normalized route weights using the softmax function: in, This represents the weights assigned to specific modality experts and selected general experts.

[0028] Each expert processes the input features separately: in, For the output of a specific modality expert, For the first The output of a general expert.

[0029] Final output features The following can be derived: This dynamic routing allows the network to emphasize evidence of a specific modality while still utilizing extensively trained experts.

[0030] (5) Step Five: Crisis Event Classification By processing mixed features using a classification head, information from different modalities can be fused, and the event category of the current sample can be identified, thereby improving the accuracy of crisis event detection. The implementation method is as follows: in, It is a fully connected layer, except for the last one. In addition to the layers, each layer has an activation function.

[0031] To optimize the classification performance of crisis event detection, we adopt the standard cross-entropy loss function: Where N is the total number of samples, and C is the number of categories. It's a real label. It is the predicted probability of category c.

[0032] The overall loss function is as follows: .

[0033] In the multi-modal crisis event classification process, the cross-entropy loss function and the text-image contrast loss function are jointly optimized to improve the classification discrimination ability and cross-modal semantic alignment effect of the model. Through the joint optimization strategy, the recognition accuracy of the model on different crisis types is significantly improved, and the model has higher robustness and generalization ability at the classification level, thereby effectively enhancing the overall performance of multi-modal crisis event detection.

[0034] It will be obvious to a person skilled in the art that the application is not limited to the details of the above-described exemplary embodiments, but that the application can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments should therefore be considered in all respects as illustrative and not restrictive, the scope of the application being defined by the appended claims rather than by the above description, and it is therefore intended that all changes and modifications that fall within the meaning and range of equivalency of the elements of the claims are to be embraced by the application. Any reference signs in the claims should not be construed as limiting the claims to the figures in which the reference signs are used.

[0035] Furthermore, it should be understood that although the present specification is described in terms of embodiments, not every embodiment according to the present specification needs to exhibit each and every characteristic specified in the specification. The specification has been described in this way in order to comply with the requirement to disclose the application in a patent application. The person skilled in the art will understand that the embodiments described in the specification can also be combined with each other in an appropriate manner to form other embodiments that can be understood by the person skilled in the art.

Claims

1. An incomplete multi-modal crisis event detection method based on memory pool and modality-aware expert system, characterized in that, Comprising the following steps: Missing modal completion: based on the text-image pairs in the training set, a cross-modal memory pool is constructed, in which the text and image features are extracted by the encoder and stored in the memory pool in vector form. For samples with missing text modal, the corresponding image features are used as query vectors to perform cross-modal retrieval based on cosine similarity in the text feature set of the memory pool. For samples with missing image modal, the text features are used as query vectors to perform cross-modal retrieval based on cosine similarity in the image feature set of the memory pool. After retrieval, the P samples with the highest cosine similarity are selected, and the Q keywords with the highest frequency are extracted from the text data in these samples. Then, for samples with missing text modal, the original image features, the P pairs of text data retrieved, and the extracted keywords are used to construct the prompt information, which is input into the image-to-text large model to generate the completed text. For samples with missing image modal, the original text features, the P pairs of text data retrieved, and the keywords are used to construct the prompt, which is input into the text-to-image large model to complete the image. Text and visual feature extraction: TEXTCNN encoder and CLIP text encoder are used to extract text features, ResNet and CLIP visual encoder are used to extract image features, and the features are concatenated by dimension, and image-text contrastive loss is used to reduce the feature difference between the completed features and the real features. Deep feature interaction: transformer-based multi-head attention mechanism is used to realize deep interaction between text and image features, extract fine-grained feature information, and mine subtle event clues in images and text. Gated expert modal perception: different experts are used to process different modal combinations to identify the modal feature differences between samples. Crisis event classification: the classification head is used to process the mixed features to identify the event category to which the current sample belongs.

2. The incomplete multi-modal crisis event detection method based on memory pool and modality perception expert system according to claim 1, wherein, In the missing modal completion step, if the missing modal is text, the image-to-text large model is used to synthesize the missing description, and if the missing modal is image, the text-to-image large model generates a visually consistent image through back diffusion under the guidance of the prompt.

3. The incomplete multi-modal crisis event detection method based on memory pool and modality perception expert system according to claim 1, wherein, In the text and visual feature extraction step, for text features, TEXTCNN encoder and pre-trained CLIP text encoder are used to extract modal-specific features and modal-consistent features: For image features, ResNET architecture and pre-trained CLIP image encoder are used to extract modal-specific features and modal-consistent features: wherein, and denote the modal-specific and modal-consistent encoders for text, respectively, and denote the modal-specific and modal-consistent encoders for images, respectively, and denote the modal-specific and modal-consistent features for text, and denote the modal-specific and modal-consistent features for images. Concatenate the modal-specific and modal-consistent features as the output features of text and image modal: wherein or , denotes a concatenation operation.

4. The memory pool and modality perception based expert system based incomplete multi-modal crisis event detection method as claimed in claim 3, wherein, The image-text contrastive loss used is as follows: where s is a learnable temperature parameter to control the smoothness of contrastive learning; denotes a similarity matrix from text features to image features; denotes a similarity matrix from image features to text features; denotes the transpose of the image matrix; denotes an exponential operation on the temperature parameter; The contrastive loss of each sample is as follows: Wherein, B is the size of each batch of training samples; is a text-to-image contrast loss; is an image-to-text contrast loss; represents the index of the current sample; represents the index of the candidate matching sample; The final contrastive loss is: wherein, is a hyperparameter for the contrastive loss weight.

5. The memory pool and modality perception based expert system based incomplete multi-modal crisis event detection method as claimed in claim 1, wherein, In the deep feature interaction step, a hybrid serial-parallel encoder is used for fine-grained feature extraction and to capture subtle event clues. The multi-head self-attention mechanism based on the transformer is used to process the text and image features respectively, to enhance the representation of each modality, to calculate the self-attention of each attention head, and to independently enhance the text and image features: in, , and They are the first The query size, key and value embedding, For learnable parameter matrix, Indicates the input feature dimension. The output dimension is used, and then the self-attention of each attention head is calculated. wherein, represents an output matrix of the head, represents a concatenation matrix of the attention outputs of the m heads, is a final output matrix representing multi-head attention, , represents the number of heads, represents a concatenation operation, represents an output linear transformation matrix; Subsequently, a FFN network with two fully connected layers and ReLU activation function is adopted to obtain the enhanced features of the modalities: ​ wherein, denotes a multi-head attention operation; The enhanced text and image features are fused through a cross-modal fusion module to generate a mixed feature: wherein, represents an image enhancement feature, is a text enhancement feature, with is a mixed feature of image and text.

6. The memory pool and modality perception based expert system based incomplete multi-modal crisis event detection method as claimed in claim 1 wherein, Different experts are used to process different modalities, including the following steps: Adopting lightweight routing network , according to semantic content and detected modal configuration prediction adaptive routing score; complete text modal When the crisis perception routing network first calculates the score of each route : where, denote the scores of three specific modal experts , , , denote the scores of general experts , from the set of general experts, the top k experts are selected according to their routing scores, denoted as , combined with the routing scores of specific modal experts , , the normalized routing weights are calculated by the softmax function: wherein, denotes the weight assigned to the specific modality specialist and the selected generalist. Each expert processes the input features respectively: wherein, the output of a specific modality expert, the output of a first generalist expert; Final output feature is derived from the equation: 。 7. The memory pool and modality perception based expert system based incomplete multi-modal crisis event detection method as claimed in claim 5, wherein, The mixed feature is used for crisis event classification: wherein, is a fully connected layer, except for the last layer which has an activation function; A standard cross-entropy loss function is used to optimize the classification performance of crisis event detection: where N is the total number of samples, C is the number of classes, is the true label, is the predicted probability of class c. The total loss function is as follows: 。

Citation Information

Cited By

  • Medical question and answer method fusing cross-modal hybrid experts

    CN121189510A

  • A medical question and answer method fusing cross-modal mixed experts

    CN121189510B

  • Expert knowledge and confrontation prompt learning-based violence and terrorism content identification model and method

    CN121637197A

  • Violent and terrorist content recognition model and method based on expert knowledge and adversarial prompt learning

    CN121637197B